Compare commits

..
Author SHA1 Message Date
TudorandClaude Opus 5 6fc7fce948 feat(flags): put /about and /blog behind flags, dark by default
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m15s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 20s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m21s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 12s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 1m39s
Both features ship dark. Neither is reachable in an environment where
its flag is off, and every flag in this system starts off, so a deploy
of this commit makes both disappear until someone turns them on
deliberately.

Two independent flags rather than one, which makes blog-on-about-off a
reachable state. That state is the whole reason the change is larger
than four notFound() calls: the blog leans on the About page for its
author identity. The Person entity is anchored at /about#tudor, and
that URL 404s while about_page is dark, so a post published in that
state would claim an author resolving to nothing. Worse than having no
named author. Both bylines fall back to unlinked text and the
BlogPosting attributes to the publisher instead, so every combination
of the two flags renders something correct.

Gated: /about, /blog, /blog/[slug], the RSS feed, both footer links,
and the matching content-sitemap entries. A sitemap must never
advertise a URL that 404s. With both dark it emits a valid empty
urlset rather than a 404, because robots.txt names it unconditionally.

Not gated: /admin. Posts have to be writable before the blog is worth
switching on, so flagging the panel would make the flag unflippable.

getFlags takes a revalidate rather than always using the 300s
constant. Reading a flag pins the calling route to the lowest
revalidate among its fetches, and the footer links live in the root
layout, so a naive gate there would have dropped every school and
place page from a weekly cache to a 5-minute one. The layout passes
604800, the floor those routes already declare, and the build confirms
all four SSG route families still prerender. The cost is one-way
latency: pages follow a flip in minutes, footer links within a week.

The e2e journeys follow the existing paired shape from the
admission_distance flag: a lit journey and a dark one for each flag,
reading state from whether /about and /blog respond rather than from
/api/flags, which another journey asserts is not publicly reachable.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DXnXQKnPpZBBP61fBQiFkq
2026-09-08 17:12:25 +01:00
tudor eb6d918650 Merge pull request 'fix(cms): regenerate the import map so the Content field renders' (#143) from fix/payload-import-map into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 15s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 1m13s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 2m9s
Reviewed-on: #143
2026-09-02 20:12:17 +00:00
TudorandClaude Opus 5 d47ac71c47 fix(cms): regenerate the import map so the Content field renders
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m11s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 11s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m10s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 11s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 52s
Creating a post in the admin panel showed no Content editor, and saving
failed validation on the field the writer was never shown.

The admin panel does not import field components. The server hands the
client a path per field, and resolves it through the generated map at
app/(payload)/admin/importMap.js. A richText field's path is
@payloadcms/richtext-lexical/rsc#RscEntryLexicalField. The committed map
held one entry, @payloadcms/next/rsc#CollectionCards, generated before
the blog collections existed and never re-run. A path missing from the
map is not an error the panel reports: the field simply does not render,
while required is still enforced server-side on save.

next build does not regenerate the map, so the stale copy shipped in the
image and the editor was equally broken on staging and production.

Regenerated with payload generate:importmap, which adds the lexical RSC
field, cell and diff components, BlocksFeatureClient for the Callout
block, and the default toolbar features.

Two things stop it drifting again. There was no script to run, so
package.json gets generate:importmap. And a test asserts the map carries
an entry for each thing the config asks for, in the source-reading style
of the other payload suites; against the old map all five fail.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DXnXQKnPpZBBP61fBQiFkq
2026-09-02 20:44:05 +01:00
tudor 17e5371e9c Merge pull request 'fix(cms): ship the initial Payload migration (recovers two commits stranded after #140 merged)' (#141) from fix/payload-initial-migration into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 1m13s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 2m13s
Reviewed-on: #141
2026-09-02 17:52:17 +00:00
TudorandClaude Opus 5 e2c63a9905 fix(cms): ship the initial migration so a container finds its tables
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m14s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 11s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m14s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 1m33s
Staging failed on boot with 42P01, relation "payload.users" does not
exist. The schema was empty because no migration existed, and the
adapter cannot create tables itself: db-postgres/connect.js gates push
on NODE_ENV !== 'production', so it is inert in a deployed container
regardless of config.

The generated migration is schema-qualified to "payload" throughout but
does not create that schema — schemaName says where tables go, it does
not create anything. It only worked against the throwaway database used
to generate it because the schema was created there by hand, so every
real environment would have failed on the first statement. CREATE SCHEMA
IF NOT EXISTS is hand-added at the top of up(), which makes it exactly
the kind of edit a regeneration discards silently; a test asserts it is
present and ordered before the first CREATE TABLE.

payload-types.ts is now committed rather than ignored. Ignoring it meant
CI typechecked against looser types than a developer with a generated
copy, which is how a Record<string, unknown> cast passed CI and then
failed locally the moment the file appeared. The post page uses the
generated Post and Media types instead, and narrows heroImage rather
than asserting it, since the field is an id at shallow depth.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017YmbBhr8s7GusjDE12hrZM
2026-09-02 18:33:27 +01:00
TudorandClaude Opus 5 3f3c5953f6 style(copy): remove em dashes from the site's prose
The em dash is one of the clearest tells of machine-written text, which
is the exact impression this work exists to remove. Rewritten rather
than substituted: where a dash was carrying a real aside the sentence is
split or recast, not patched with a comma.

Covers the About page, the two Callout labels an editor sees in the
admin panel, and PUBLISHING.md, which defines the house style and should
follow it. The rule is now recorded in that house style and in the
spec's voice rules, so it survives this branch.

Code comments are left alone: they are not copy, and the surrounding
codebase uses the same punctuation throughout.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017YmbBhr8s7GusjDE12hrZM
2026-09-02 18:33:27 +01:00
tudor 124c6702a9 Merge pull request 'feat: a named author, an About page and a Payload blog' (#140) from feat/about-and-blog into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 45s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 1m32s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 2m8s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 2m36s
Reviewed-on: #140
2026-09-02 16:04:58 +00:00
TudorandClaude Opus 5 e25722d9ab fix(blog): hide drafts at the access layer, and back the --drop claim
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m12s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 32s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m9s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 1m15s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 2m26s
Review findings on #140.

Drafts were reachable. Posts granted unconditional public read and the
_status filter lived only in the pages that query the collection — which
is a convenience, not a control. Payload's documentation is explicit:
"The `draft` argument alone does not restrict documents with _status:
'draft' from being returned by the API." A direct GET /cms-api/posts
would have handed every unpublished draft to any visitor. Read access
now returns a query constraint for anonymous callers, which is the
documented mechanism.

The --drop claim was asserted across four files while the spec still
listed it as an open question. Now verified rather than assumed:
run_full_migration drops exactly ["school_results", "schools"] by name,
there is no drop_all() or DROP SCHEMA anywhere in backend/, the only
other drop is schema-qualified to marts, and nothing sets search_path.
The guarantee is stronger than schema isolation alone — those two table
names do not exist in Payload — so the claim stands, but it now rests on
cited code. The spec records the evidence and closes the open item.

findPost is wrapped in React's cache(): Next calls generateMetadata and
the page separately for one request, so every post view ran the same
query against Postgres twice.

The bare .lede rule was dead — .prose p scores (0,1,1) and outranks it —
so only .prose .lede ever applied. Removed, with the specificity noted
so the surviving selector is not "simplified" back into a silent
regression.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017YmbBhr8s7GusjDE12hrZM
2026-09-02 16:48:59 +01:00
TudorandClaude Opus 5 07d586d0ad docs(blog): how to publish, and why the app has two route groups
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m15s
PR Checks / Backend Smoke (pull_request) Successful in 10s
PR Checks / Build Backend (no push) (pull_request) Successful in 36s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m11s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 1m18s
PR Checks / AI Code Review (Claude) (pull_request) Failing after 2m43s
PUBLISHING.md carries the house style with the posts, so the standard
survives without the design doc to hand — including the rule that a post
states what a metric does not show, which is the strongest signal a
human wrote it.

CLAUDE.md gains the two constraints that are invisible from the code and
expensive to rediscover: metadata file conventions break if moved into a
route group, and the build must keep succeeding with DATABASE_URL unset.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017YmbBhr8s7GusjDE12hrZM
2026-09-02 16:30:09 +01:00
TudorandClaude Opus 5 b793640507 feat(blog): add the blog index, post pages, RSS and content sitemap
The rendering split is dictated by CI building with no database.

/blog, /blog/rss.xml and /content-sitemap.xml have no dynamic params, so
Next prerenders them at build time and the build fails on a missing
Payload secret — caught here, not on staging. They are force-dynamic
instead: one indexed query against Postgres on the same Docker network,
and a newly published post appears immediately rather than waiting on a
revalidation. /blog/[slug] keeps ISR, because with no
generateStaticParams there is nothing to prerender; it is generated on
first request and cached, which is exactly what the collection's
afterChange hook exists to invalidate.

RichText takes `converters`, not `blocks`, in Payload 3.88, and the
default converters must be spread or every paragraph and heading loses
its renderer and the body comes out empty.

BlogPosting references the Person and Organization by @id rather than
repeating them, so every post and the About page resolve to one author
entity instead of declaring several people with the same name.

/sitemap.xml is proxied from FastAPI, which knows nothing about Payload,
so the Next-owned URLs get their own sitemap and robots.txt lists both.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017YmbBhr8s7GusjDE12hrZM
2026-09-02 16:28:58 +01:00
TudorandClaude Opus 5 21a5d18f59 feat(blog): add the posts and media collections
Drafts are on so a post can be written across sittings without saving
being publishing.

afterChange and afterDelete revalidate every path a post appears on.
Blog pages are ISR because CI builds with no database, so without these
a published post would not appear until the revalidate window expired —
up to an hour of a writer concluding that publishing is broken. Payload
runs in the same process as Next, so these are direct revalidatePath
calls with no webhook and no shared secret.

Media writes to an absolute /app/media matching the compose mount; a
mismatch would write into the container filesystem, where the next
redeploy silently discards it. Alt text is required rather than
optional.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017YmbBhr8s7GusjDE12hrZM
2026-09-02 16:24:15 +01:00
TudorandClaude Opus 5 f614414070 feat(about): give the site a named author
The site had no author, no statement of why it exists and nobody
accountable for its numbers, which is most of why it reads as machine
generated.

The page states plainly that its author is not an education expert. The
credibility claim is lived experience — a parent going through primary
admissions — plus stated provenance for every figure, which is true and
cannot be undermined by someone noticing there is no teaching
qualification behind it. First name only: the Person JSON-LD carries no
familyName, worksFor or affiliation, and a test asserts it stays that
way.

The footer gains a fourth column, with a tablet breakpoint so four
columns pair up rather than crushing before the 768px collapse. The nav
is deliberately untouched — its mobile tab bar already carries four
items.

public/brand/tudor.jpg is NOT in this commit. The page references it and
will show a broken image until the photograph is supplied; a stock
portrait would defeat the entire point of the work.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017YmbBhr8s7GusjDE12hrZM
2026-09-02 16:22:52 +01:00
TudorandClaude Opus 5 310b63b0cb build(cms): wire Payload into the Docker image and both stacks
Uploads go to a named volume at /app/media. The directory is created in
the image before the mount and covered by the existing chown, because
Docker seeds a fresh named volume from the image path — a missing or
root-owned directory there fails every upload with EACCES at runtime,
long after the build passed.

PAYLOAD_SECRET uses the same :? form as AIRFLOW_ADMIN_PASSWORD: refuse
to start rather than boot with an empty secret and accept forged
sessions. Staging's must differ from production's, which the header
comment now says explicitly. Portainer prefixes volume names per stack,
so payload_media isolates itself.

prodMigrations is not wired yet — generating the initial migration needs
a reachable Postgres. Follows in its own commit.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017YmbBhr8s7GusjDE12hrZM
2026-09-02 16:19:27 +01:00
TudorandClaude Opus 5 c2c76c5817 feat(cms): keep the admin panel out of the index
X-Robots-Tag rather than the robots.txt Disallow alone, for the same
reason the staging rule uses one: a Disallow blocks crawling, not
indexing, so a URL found from an external link can be indexed without
ever being fetched — and blocking the crawl means the noindex is never
seen. Both mechanisms are applied to /admin and /cms-api.

The existing CSP is frame-ancestors only, which restricts who may embed
the site rather than what a page may load, so it cannot break the panel.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017YmbBhr8s7GusjDE12hrZM
2026-09-02 16:17:55 +01:00
TudorandClaude Opus 5 c5a4d106da feat(cms): install Payload and serve the admin panel
Payload 3.88 runs inside the Next app against the existing Postgres, in
its own 'payload' schema so no pipeline operation on public — the app
tables, Airflow's metadata, migrate_csv_to_db.py --drop — can reach blog
content.

Its REST API is mounted at /cms-api. /api is the FastAPI proxy's
catch-all, which would swallow every admin call and forward it to the
backend with no error. The mount points live in lib/payloadRoutes.ts so
there is one definition and a test can assert it without importing
Payload: it is ESM-only, next/jest will not transform it, and appending
transformIgnorePatterns cannot un-ignore a package. Forcing it through
transpilePackages would change how the production build bundles Payload
to serve a test, so the live proof that /api still reaches FastAPI stays
where it belongs — the e2e journeys, which call /api/schools.

The package becomes ESM ("type": "module"), which Payload's CLI requires:
richtext-lexical has top-level await and the config cannot be require()d.
Only two files needed renaming, jest.config.cjs and a build script.

The build is verified to succeed with DATABASE_URL and PAYLOAD_SECRET
both unset, which is how CI builds it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017YmbBhr8s7GusjDE12hrZM
2026-09-02 16:17:01 +01:00
TudorandClaude Opus 5 2437ffce42 refactor(app): move site routes into a (frontend) route group
Payload's admin panel ships its own root layout rendering html/body.
Next allows multiple root layouts only when no app/layout.tsx exists, so
the site's routes move into their own group. Route groups are invisible
to routing: every public URL is unchanged, verified against the build's
route table.

The metadata file conventions deliberately stay at the app/ root. Moving
them into the group renamed /icon.png to /icon-4usi79.png (likewise
apple-icon and opengraph-image) and dropped /robots.txt altogether,
which would have broken the /icon.png cache-control rule, the
outputFileTracingIncludes entry for the share card, and robots.txt.

darkThemeSafety reads app/globals.css off disk rather than importing it,
so it needed its own path fix — a grep for import specifiers misses it,
and it fails as an unrunnable suite rather than a failed assertion.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017YmbBhr8s7GusjDE12hrZM
2026-09-02 16:11:48 +01:00
TudorandClaude Opus 5 eb648f3f76 build(next): convert the config to ESM so Payload can wrap it
withPayload() is ESM-only, so next.config.js has to become .mjs. That
file also carries the rule that keeps staging out of Google's index, so
the conversion goes in on its own, behind a test that asserts the rule
survived — along with the standalone output, the opengraph-image font
tracing and the analytics frame-ancestors CSP.

Jest resolves the .mjs config without extra configuration, so
jest.config.js is untouched.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017YmbBhr8s7GusjDE12hrZM
2026-09-02 16:08:58 +01:00
TudorandClaude Opus 5 74e5fffc10 docs(about-blog): implementation plan for the About page and Payload blog
Nine tasks, each ending in an independently testable deliverable.

Two structural findings that the spec did not anticipate, both recorded
in the plan. Payload's admin panel ships its own root layout rendering
html/body, and Next allows multiple root layouts only when no
app/layout.tsx exists — so every existing route moves into an
app/(frontend) route group first, on its own, with the full suite as the
gate. Route groups are invisible to routing, so no public URL changes.

The second finding corrects the spec: adding /cms-api to the FastAPI
proxy's exclusion list would be dead code, because that catch-all only
ever matches /api/*. The route remap alone is sufficient.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017YmbBhr8s7GusjDE12hrZM
2026-09-02 15:59:44 +01:00
TudorandClaude Opus 5 748ef32180 docs(about-blog): design for a named author, an About page and a Payload blog
The site reads as synthetic because nobody is accountable for the
numbers, no editorial judgement is visible, and the voice is
institutional third person. This designs the fix: a named author
(first name, photo, explicitly not an education expert), a coded
/about page, and a blog backed by Payload CMS running inside the
existing Next app.

Also fills a hole in the SEO programme, which has eight workstreams
and no E-E-A-T or authorship signal on a YMYL corpus.

Records two collisions found while designing, both of which fail
badly if missed: Payload's default /api route fights the existing
FastAPI catch-all proxy, and withPayload() is ESM-only so
next.config.js — which carries staging's noindex header — has to
become next.config.mjs.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FT1Ls4GbgLDXoQX7NAuHGT
2026-09-02 14:43:20 +01:00
tudor b0c4ea8282 Merge pull request 'fix(destinations): the DAG died on a null in a primary key' (#139) from fix/destinations-national-grain into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 14s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 51s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 1m36s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Successful in 2m6s
Reviewed-on: #139
2026-08-31 20:43:26 +00:00
TudorandClaude Opus 5 e236669fde fix(destinations): school rows and the England reference are different grains
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m4s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 11s
PR Checks / Build Frontend (no push) (pull_request) Successful in 45s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 52s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 2m27s
The annual DAG died with a BrokenPipeError from Meltano's log writer, which
is several frames from the cause: target-postgres exited first and the tap
saw its stdout close.

The tap declared primary_keys = [urn, ...] while emitting urn=None for the
national rows, and target-postgres turns primary_keys into a NOT NULL
constraint. The first national row of the run failed the insert and took
the loader with it. Every other tap in this repo keys on non-null columns.

Carrying two grains in one stream was the actual mistake, so the fix is to
separate them rather than paper over the null: four streams now, with
ees_ks4/ks5_destinations_national carrying no urn column at all — a school
identifier that is null in every row is a grain mismatch, not a column.
The staging models split the same way and the national mart reads the new
pair instead of filtering `where urn is null`.

Verified against the live API: the school stream yields 135,240 rows over
4,508 schools with no duplicate keys, no null key columns and all 31,382
suppression sentinels intact; the national streams yield 30 and 33 rows
with no urn column.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvdDKvFFSZuMVDH5fEyTob
2026-08-31 21:36:43 +01:00
tudor fb5a0928bd Merge pull request 'fix(airflow): a fixed admin password from the environment, and a tap that missed #137' (#138) from fix/airflow-fixed-admin-password into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 55s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 1m36s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Successful in 2m14s
Reviewed-on: #138
2026-08-31 19:50:00 +00:00
TudorandClaude Opus 5 264edd2e3a fix(airflow): a login that survives a container restart
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m7s
PR Checks / Backend Smoke (pull_request) Successful in 10s
PR Checks / Build Backend (no push) (pull_request) Successful in 12s
PR Checks / Build Frontend (no push) (pull_request) Successful in 50s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 51s
PR Checks / AI Code Review (Claude) (pull_request) Failing after 2m58s
The simple auth manager generates a random password on first start and
writes it to a file, so every restart of the api-server invalidated the
last one and the password had to be dug out of the container logs again.

The stack now writes that file itself from AIRFLOW_ADMIN_PASSWORD before
exec'ing the api-server. Airflow generates nothing when the file already
exists, so the login is whatever the stack environment says it is.

Written with python rather than echo, so json.dumps escapes a password
containing quotes, backslashes or non-ASCII correctly — verified against
`p@ss "wo\rd' £5`, which round-trips intact.

An unset AIRFLOW_ADMIN_PASSWORD raises KeyError and the container exits.
Falling back to a generated password would silently undo the point of the
change, and a compose-level `:?` gives the same refusal a readable reason.
This does mean the variable MUST be set in Portainer before the next
deploy of either stack.

Not affected by the two Docker gotchas in the upstream docs: this image
has no USER directive so it runs as root, and the file is rewritten from
the environment on every start rather than persisted on a volume.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvdDKvFFSZuMVDH5fEyTob
2026-08-31 17:38:42 +01:00
TudorandClaude Opus 5 cd2cbe7be6 build(pipeline): install the destinations tap in the image
meltano install would resolve it from pip_url, but five of the six custom
taps are also installed explicitly and a new plugin failing to appear is
not something you want to debug from a deploy log.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvdDKvFFSZuMVDH5fEyTob
2026-08-31 17:37:46 +01:00
tudor 73182d0c0c Merge pull request 'feat(destinations): say what happened to a school's leavers, without republishing what DfE withheld' (#137) from feat/ks4-destinations into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 20s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 50s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 2m6s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Successful in 2m2s
Reviewed-on: #137
2026-08-30 20:49:14 +00:00
TudorandClaude Opus 5 cbe3a9a772 fix(destinations): the table said 'withheld' for a category that just doesn't apply
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 18s
PR Checks / Build Frontend (no push) (pull_request) Successful in 45s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 1m14s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 9s
The Share column keyed off `percentage === null`, which is true for
not_applicable as well as suppressed, so a destination that does not apply
to the school was labelled as one DfE withheld — while the Pupils column
in the same row rendered blank. Two columns, one row, disagreeing about
what the row was, and one of them making a claim about DfE that wasn't
true.

Both columns now derive from `status`, which is the distinction the mart,
the SQLAlchemy model and the serialiser all preserve deliberately:
published shows the figure, suppressed shows the withheld badge,
not_applicable shows an em-dash with a title saying so.

A published count with no published percentage now derives its share from
the cohort rather than falling through to a marker — both halves are
published, so nothing withheld is involved, and it is the same derivation
the bar widths already use.

Verified the new tests fail against the old logic before keeping them.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvdDKvFFSZuMVDH5fEyTob
2026-08-30 21:48:39 +01:00
TudorandClaude Opus 5 2e9b5c83c5 fix(destinations): the masking pass can no longer exit unsafely
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m5s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 18s
PR Checks / Build Frontend (no push) (pull_request) Successful in 45s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 1m16s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 7m23s
Review found _mask_for_disclosure could return with its invariant broken
and say nothing. add_companion only ever withheld a *published* cell, so a
group with one suppressed category and every other one not_applicable —
routine in special schools and AP, where few categories apply — left the
loop with the lone suppressed cell still solvable. Reproduced on a
nine-pupil cohort: one hidden cell, cohort served, residual intact.

A disclosure-control pass that fails silently is worse than none, because
everything downstream trusts it. The loop now runs until the invariant
holds and escalates when no companion exists: the pupil group is dropped
from the payload, and an empty block serialises as None so the section is
absent rather than an empty shell. disclosure_invariant_holds() is exported
so tests assert it directly instead of re-deriving it, and an exhaustive
test sweeps all 81 suppression patterns of a four-category group.

Also fixes a test that set up six measures and checked one: the loop was
`for measure in ["school_sixth_form"]`. It now checks every measure, and
against the real invariant — none hidden, or at least two, rather than
"at least two", which the five published measures would have failed.

No regression on real data: 262 mainstream secondaries, all-pupils bar
still drawable on 94%, zero invariant violations, one disadvantaged group
dropped by the new escalation.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvdDKvFFSZuMVDH5fEyTob
2026-08-30 21:35:03 +01:00
TudorandClaude Opus 5 102397fe69 fix(destinations): withhold at the API, not just in the chart
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 17s
PR Checks / Build Frontend (no push) (pull_request) Successful in 45s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 1m14s
PR Checks / AI Code Review (Claude) (pull_request) Failing after 3m49s
Code review found the disclosure the whole design was meant to prevent.
R1 was written as a rendering rule and implemented as one: canRenderBar
stopped the bar being drawn, but GET /api/schools/{urn} still carried the
cohort and every published category. cohort - sum(published) returned
Whitley Bay's withheld further-education figure exactly — 18 pupils — to
any caller, and the RSC payload put it in the browser too.

app.py already stated the principle for admission_distance: this endpoint
is public and unauthenticated, so a field left in the payload is a
published field. The same reasoning applies here and did not get applied.

_mask_for_disclosure now closes both identities before serialisation —
categories sum to the cohort, and disadvantaged + other = all — by adding
secondary suppression until every row and column hides none or at least
two. My first attempt picked the smallest published cell as the companion
and a new test caught it choosing a zero, which protects nothing: the
residual still resolved to 18. The companion must carry pupils.

DfE's own aggregates are no longer served. Nothing rendered them, and one
spanning a single suppressed component names it.

Cost, measured over 262 mainstream secondaries: the all-pupils bar
survives on 94% rather than 100%. Zero lone-suppressed groups remain.

The e2e helper now tells a missing feature apart from missing data: it
fails if the API serves no destinations key at all, and skips if the key
is served but the annual DAG has not populated the marts. Failing on the
second would redden the staging gate for unrelated commits.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvdDKvFFSZuMVDH5fEyTob
2026-08-28 19:58:02 +01:00
Tudor 68a192e430 Merge remote-tracking branch 'origin/main' into feat/ks4-destinations
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 16s
PR Checks / Build Frontend (no push) (pull_request) Successful in 45s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 1m12s
PR Checks / AI Code Review (Claude) (pull_request) Failing after 4m4s
# Conflicts:
#	nextjs-app/__tests__/components/darkThemeSafety.test.ts
2026-08-28 18:45:27 +01:00
TudorandClaude Opus 5 ccd5074c90 test(e2e): destination journeys, including the no-bar rule
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m5s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 19s
PR Checks / Build Frontend (no push) (pull_request) Successful in 47s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 1m16s
PR Checks / AI Code Review (Claude) (pull_request) Canceled after 1m10s
The helper throws rather than skipping when no school returns a
destinations block: a silent skip would let a real regression in the
sections ride along unnoticed, which is why the distance journeys were
changed the same way in 4f01fbd.

The disadvantaged journey computes the residual itself and asserts it
appears nowhere on the page — the one number the section must never state.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvdDKvFFSZuMVDH5fEyTob
2026-08-28 16:16:25 +01:00
TudorandClaude Opus 5 2b4cf20d75 feat(destinations): the post-16 section, replacing the placeholder
The 'Post-16 destination data coming soon' note is deleted rather than
reworded: for a school with no sixth form the truthful statement is that
the question does not apply, and a placeholder there implies something is
missing. hasSixthForm and .sixthFormNote go with it — nothing else used them.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvdDKvFFSZuMVDH5fEyTob
2026-08-28 16:15:21 +01:00
TudorandClaude Opus 5 cef2f77149 feat(destinations): the After Year 11 section
Question cards over one bar, with the cards acting as a lens on the bar
rather than a summary beside it — focusing a card dims everything it is
not made of, so the grouping we chose is inspectable rather than asserted.

The bar renders only when canRenderBar allows it. Where a category is
withheld the section says so and shows the table instead: the categories
sum to the cohort, so a bar drawn from the published segments leaves a
gap whose width is the withheld figure.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvdDKvFFSZuMVDH5fEyTob
2026-08-28 16:14:13 +01:00
TudorandClaude Opus 5 68b6417149 feat(destinations): types and secondary section flags
Destinations are secondary-only, so the flags go on computeSecondaryFlags
rather than computeSchoolFlags. A phase counts as present only when some
pupil group carries categories — an empty block would otherwise open a nav
entry pointing at a section that never renders.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvdDKvFFSZuMVDH5fEyTob
2026-08-28 16:11:37 +01:00
TudorandClaude Opus 5 c5719ef362 feat(destinations): serve destinations without closing the gaps
The serialiser carries status through and computes no totals of its own.
The only aggregates in the payload are ones DfE published itself; whether
showing one is safe depends on how many of its components are suppressed,
which the frontend decides.

The batch guard grows from six tables to eight.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvdDKvFFSZuMVDH5fEyTob
2026-08-28 16:10:29 +01:00
TudorandClaude Opus 5 5e5b61987a feat(destinations): marts, with R3 masking applied at the boundary
Disadvantaged and other-pupils partition the whole and the all-pupils
figure is published, so publishing both halves recovers the suppressed
one. The mask is applied in the mart rather than the API so no consumer
added later can reach an unmasked combination.

The R1 test is a warn, not an error: DfE publishes the recoverable
combination and the mart's job is to carry it faithfully. Refusing to
close the gap is the API's job and the frontend's.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvdDKvFFSZuMVDH5fEyTob
2026-08-28 16:08:20 +01:00
TudorandClaude Opus 5 c564566432 feat(destinations): staging models that keep 'withheld' distinct from 'absent'
safe_numeric maps every EES sentinel to NULL, which is right for attainment
and wrong here: one of those states has to print 'withheld' and the other
has to print nothing. A status column carries the difference.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvdDKvFFSZuMVDH5fEyTob
2026-08-28 16:07:17 +01:00
TudorandClaude Opus 5 9188626051 feat(destinations): a tap that preserves the suppression sentinel
EES writes 'c' where a figure is withheld and the categories sum to the
cohort, so counts and percentages are emitted as text with the sentinel
intact. safe_numeric must never be pointed at them.

School rows and the England reference need different establishment pins:
at national level selective schools, studios and UTCs are separate
populations rather than labels, so leaving establishment open multiplies
30 rows into 190. Two queries per period, each keeping its own level.

Verified against the live API for 2022/23: 135,240 school records over
4,508 schools, exactly 30 each, no duplicate keys, 31,382 sentinels kept.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvdDKvFFSZuMVDH5fEyTob
2026-08-28 16:06:37 +01:00
TudorandClaude Opus 5 7ae9ecdc36 feat(destinations): colour tokens, with the absence hatched not coloured
Activity not captured includes independent schools and moving abroad, so a
red segment would be a factual error. The hatch doubles as the secondary
encoding that rescues the neutral/blue pair, which separates at only dE 7.6
as flat fills.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvdDKvFFSZuMVDH5fEyTob
2026-08-28 15:58:25 +01:00
TudorandClaude Opus 5 1980d79eee feat(destinations): the disclosure rules, as executable guards
The destination categories sum to the cohort and DfE publishes the cohort
total, so a lone suppressed cell is recoverable by subtraction. canAggregate,
canRenderPublishedAggregate and canRenderBar are what stop a consumer doing
that; toBarSegments throws rather than leaving a readable gap.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvdDKvFFSZuMVDH5fEyTob
2026-08-28 15:57:46 +01:00
TudorandClaude Opus 5 9423f11567 docs(destinations): implementation plan, ten tasks
Ordered so the disclosure guards land first and everything downstream
consumes them: lib/destinations.ts, tokens, tap, staging, marts, API,
then the two sections and the journeys.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvdDKvFFSZuMVDH5fEyTob
2026-08-28 15:56:36 +01:00
TudorandClaude Opus 5 576013d627 docs(destinations): design for KS4 and post-16 destination measures
The published files suppress individual cells, not whole cohorts, and the
categories sum to the cohort — so on 22% of mainstream secondaries the
withheld figure can be recovered by subtraction. Three disclosure rules
fall out of that, and the rest of the design is downstream of them.

Verified against the EES API rather than assumed: both datasets carry
school-level rows keyed by URN, with the disadvantage split and every
destination category the display needs.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvdDKvFFSZuMVDH5fEyTob
2026-08-28 15:28:09 +01:00
tudor 7c08138fe4 Merge pull request 'fix(map): the popup never took the dark theme' (#136) from fix/dark-mode-map-popup-contrast into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 49s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Successful in 1m41s
Reviewed-on: #136
2026-08-27 21:57:53 +00:00
TudorandClaude Opus 5 a7829d591a fix(map): the popup never took the dark theme
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 10s
PR Checks / Build Frontend (no push) (pull_request) Successful in 46s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 15s
leaflet.css paints `background: white; color: #333` on the popup card and its
tip. LeafletMapInner binds themed content into it — the school name and the
headline figure are var(--text-primary) — so in dark mode #E9EEF0 landed on
#FFFFFF at 1.17:1. The two things the popup exists to say were the two least
readable things on the page.

Every other foreground in that popup failed too, from the same cause: the
muted phase line at 2.90:1, the vs-national delta at 1.94:1, the Ofsted badge
at 1.74:1. Moving the surface onto --bg-card fixes all of them at once —
13.52, 5.45, 8.14 and 9.11:1 respectively. In light mode --bg-card is #FFFFFF,
so the popup renders exactly as it did.

globals.css already pulls the rest of Leaflet's chrome onto the tokens, and
says why: "this matters most in dark mode, where Leaflet's white attribution
bar would otherwise sit on a near-black page." The popup was simply missed.

The View Details button needed its own fix. It pairs background:var(--status-
above) with a literal white label, which theming the card does not reach:
--status-above is #36743F in light but #7FCB8A in dark, taking the label from
5.63:1 to 1.94:1. --text-inverse is the token for ink on a saturated fill, and
the popup's own Ofsted badge already uses it.

darkThemeSafety already guards this defect class, but only inside .module.css.
Neither half of this one lives there — the surface is a third party's, the
text is inline in a TSX template — so it scanned clean throughout. Two rules
added for the layer it could not see. Fixing the grouped-selector blind spot
in its rules() helper was needed to write them: taking only a selector's last
line discarded every selector in a grouped rule but the final one, which makes
a safety guard fail open.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FuPUioHpxtaiDNagQvjxyM
2026-08-27 22:34:18 +01:00
tudor 1ed4470fc2 Merge pull request 'fix(admissions): flag-off pages must not speak for the council' (#135) from fix/distance-flag-off-absence-copy into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 52s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Successful in 1m43s
Reviewed-on: #135
2026-08-27 20:30:25 +00:00
TudorandClaude Opus 5 7a16b1b52f fix(admissions): flag-off pages must not speak for the council
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m4s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 11s
PR Checks / Build Frontend (no push) (pull_request) Successful in 46s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 1m8s
The secondary admissions section words the absence of a cut-off distance:
"<LA> has not published a cut-off distance for this school." That sentence
is true when the authority publishes nothing. It is false when the
authority does publish and the admission_distance flag is simply off — and
off is the current state, so every secondary page with an EES admissions
row has been making a claim about a council on our behalf.

The backend already draws the distinction the copy needs. /api/schools/{urn}
omits the admission_distance key entirely while the flag is dark rather than
sending null, precisely so that "we are not publishing cut-offs" stays
distinguishable from "this school has no cut-off"; lib/types.ts says so in
as many words. The page then collapsed the two with `?? null` before the
section ever saw them.

So stop collapsing it: thread the raw field to SecondarySchoolSections and
word the absence only when the feature is on. Null still gets the sentence
naming the authority — that case is unchanged and still tested.

Primary pages are unaffected: AdmissionsSection carries no absence copy and
renders nothing when there is no figure. DistanceSection already treated
absent and null alike; only its type widens.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FuPUioHpxtaiDNagQvjxyM
2026-08-27 21:17:16 +01:00
tudor cf9d41b476 Merge pull request 'fix(analytics): the funnel source read a referrer that never changes' (#134) from fix/navigation-source-soft-nav into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 52s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Successful in 1m40s
Reviewed-on: #134
2026-08-27 08:26:45 +00:00
TudorandClaude Opus 5 e820e7fecd fix(analytics): the funnel source read a referrer that never changes
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m5s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 10s
PR Checks / Build Frontend (no push) (pull_request) Successful in 47s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 1m13s
The staging E2E gate has been red since #132 merged (run 1064, and
1066 after it): "a school reached from a location page is attributed
to it, not to direct" expects `place`, receives `direct`.

#132 fixed a real bug — `/schools/` had no case and fell through to
`direct` — but the mechanism underneath it never worked.
getNavigationSource read document.referrer, which the browser writes
only when a *document* loads. Every internal navigation here is an App
Router soft navigation: history.pushState, no new document, so
document.referrer goes on naming whatever opened the tab for the whole
session.

Verified on staging: load /schools/brentwood, click a school, the URL
becomes /school/… and document.referrer is still "".

So `from` reported `direct` for essentially every in-app journey, not
just the ones through the location layer — search, rankings, compare
and detail were all being counted as "typed the URL". The unit suite
passed throughout because every case set document.referrer directly,
which only happens on a full page load.

The fix is a module-level trail written by RouteTrail, a render-nothing
client component in the root layout. Its lifetime is exactly right: it
survives soft navigation, and it dies on a real document load — which
is precisely when document.referrer becomes meaningful again, so the
two cover each other with no overlap.

Reading it skips entries equal to the current path rather than taking
the second-to-last. That makes the answer independent of whether the
layout effect or the page effect ran first — React orders those by
tree position, which is not a contract worth resting a measurement on
— and it gives the right answer both when the user returns to a page
they came from and on a hard load of a school page.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FuPUioHpxtaiDNagQvjxyM
2026-08-27 09:22:37 +01:00
tudor 4fdeb70a93 Merge pull request 'feat(places): say what each school is, not only how it scored' (#133) from feat/place-school-attributes into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 19s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 50s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 1m42s
Reviewed-on: #133
2026-08-27 07:59:10 +00:00
TudorandClaude Opus 5 9a1f56c431 feat(places): say what each school is, not only how it scored
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m4s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 17s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 2m44s
The location tables carried one column: a percentage. A parent
shortlisting from a town page is asking a different question first —
does it take my child's age, is it a faith school, does it have a
nursery — and the page could not answer any of it.

Primary tables gain Ages, Religious character, Nursery and
Constituency; secondary tables the same minus Nursery, which is a
question about a different intake. An all-through school renders in
both groups, so its nursery shows under primary alone.

The measure moves to the second column rather than the last. Six
columns overflow a phone and .tableWrap turns that into a horizontal
swipe; with the measure last, the one number the page exists for is
the one scrolled off the screen.

Cell rules are the ones the school page already uses, so the two
surfaces cannot disagree about the same school: "Does not apply",
"None" and "Not applicable" all read as no religious character, and
the en-dash age normalisation moves into formatAgeSpan, which
formatAgeRange now delegates to.

Backend: nursery_provision and parliamentary_constituency were not in
the place response. Both are optional GIAS mart columns that
data_loader degrades to NULL, and the `in rows.columns` guard keeps a
mart the pipeline has not rebuilt working.

Also fixes a live bug on the same line: SCHOOL_COLUMNS already ends
with latitude and longitude, and the endpoint concatenated them again,
so pandas dropped one of every duplicated pair and warned "columns are
not unique" on each request. Ordered de-duplication removes the
warning and the silent drop.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FuPUioHpxtaiDNagQvjxyM
2026-08-27 08:49:14 +01:00
tudor ade9dbb3ba Merge pull request 'feat(analytics): measure the location layer, and stop calling it direct' (#132) from feat/place-analytics into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 51s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 1m44s
2026-08-27 07:29:48 +00:00
TudorandClaude Opus 5 d1a8596208 feat(analytics): measure the location layer, and stop calling it direct
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 11s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 11s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 40s
The location pages were only half-tracked. Umami counts a pageview for
each of the ~3,900 URLs automatically, but nothing else: components/
places contained no track() call, and place_viewed was not even a
declared event name.

The part that mattered was worse than a gap. getNavigationSource mapped
a same-origin referrer to a funnel source and had no case for /schools/,
so every school view arriving through the location layer fell through to
'direct' — the bucket you read as "typed the URL, no referrer". W2's
whole purpose is funnelling search traffic onto school pages, so the one
measurement that says whether it worked was reporting the wrong answer,
and reporting it confidently. Verified live against staging: expected
"place", received "direct".

/schools/ is checked before /school/. They differ by one letter and mean
different things — the location layer versus a single school — and a
prefix test in the wrong order silently merges them.

place_viewed carries kind, slug, phase and school_count. kind is the
reason it exists: whether to keep investing in these pages turns on
which sort earns engagement, and a pageview cannot say, because all four
families share the /schools/ prefix and only the registry knows which is
which. It is a client component because PlaceView is a server component;
one line in PlaceView covers all four families, since they all render
through it.

Both E2E journeys were verified failing against staging first — one
because place_viewed does not exist there, the other on the exact
"place" vs "direct" mismatch.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-27 08:23:21 +01:00
tudor a3c09d9b67 Merge pull request 'fix(suggest): the dropdown reopened on top of the search results' (#131) from fix/suggest-reopens-over-results into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 50s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Successful in 1m37s
Reviewed-on: #131
2026-08-26 21:07:55 +00:00
tudor a7f4c86464 Merge pull request 'fix(search): the mobile hero search was indented by a card's padding' (#130) from fix/mobile-hero-search into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 50s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 5m50s
Reviewed-on: #130
2026-08-26 20:51:20 +00:00
TudorandClaude Opus 5 0804566736 fix(test): drop a committed scratch probe, and close a hole in the guard
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 11s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 9s
Code review, all three findings valid.

e2e/tests/__m.spec.ts was a throwaway probe used to measure the mobile
hero geometry. It asserts nothing, so it could never fail; it carried a
leftover `pick('form').constructor === Object ? null : null` that is
null either way and throws if no form matches; and it should never have
been committed. Deleted.

It survived because `rm -f e2e/tests/__m.spec.ts` ran with the shell
already inside e2e/, so the path resolved to e2e/e2e/tests/... — which
does not exist, and rm -f is silent about that. `git add -A` then swept
it in. I checked `git diff --stat` before committing, which lists only
tracked modifications and never shows an untracked file; `git status
--short` would have.

The scoping guard compared the last line of a rule's prelude against the
literal '.filterBar', so a regression written as a selector list —
`.filterBar, .other { padding }`, or the same split across two lines —
would have walked straight past the test meant to catch it. Selectors
are now split on commas and matched individually, and comments are
stripped first so a brace inside one cannot desynchronise the parse.

Verified against all three shapes: bare, inline comma list, and
multi-line comma list. Each is caught; each passes again once reverted.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-26 21:49:48 +01:00
TudorandClaude Opus 5 55363cbd18 fix(suggest): the dropdown reopened on top of the search results
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m2s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 11s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 55s
Three staging-gate failures, two of them one real bug.

After a search, the results-page bar still holds the term in its input,
so on every render the query was >= 2 characters and the suggestion list
opened again — on top of the very results the search had just produced.
Playwright reported it as "<li role=option ...> intercepts pointer
events" while trying to click the first result; a reader would simply
have found their first result unclickable. Both the school-detail and
hero-map journeys failed on it, and neither is about autosuggest.

Suggestions now answer typing, not the mere presence of a value:
`hasTyped` gates the hook, is set on change, and is cleared when a
search is submitted or a suggestion is chosen. A pre-filled input makes
no request and shows no list.

Third failure was my test, not the product. An unphased place page
renders one table per phase, and an all-through school legitimately
appears in both — so the page's school links were never one alphabetical
run. The assertion collected them all together and only passed because
no town it picked had held an all-through school. When the data gave
Abbots Langley one, Breakspeare School appeared in the primary table and
again in the secondary, and the test failed on correct behaviour. It now
checks each table separately, and passes against the data that broke it.

Guards: a jest test that a pre-filled input neither fetches nor opens
(verified by reverting — it is the only one that fails), and an E2E
journey that submits a search and then requires the first result to be
clickable, which is the reader-facing version of the same thing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-26 21:38:16 +01:00
tudor 868eb344f5 Merge pull request 'fix(map): the hero map's fade to the header was hardcoded white' (#129) from fix/dark-map-fade into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 51s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 0s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 5m49s
Reviewed-on: #129
2026-08-26 20:32:45 +00:00
TudorandClaude Opus 5 0b15497c09 fix(search): the mobile hero search was indented by a card's padding
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 11s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 1m3s
Measured at 390px: the headline and lede sit at x=34, while the search
box, the hint and the location link all sat at x=48 and the field was
28px narrower than the copy above it.

The 14px came from `@media (max-width: 768px) { .filterBar { padding:
0.875rem } }`. That rule is for the results filter bar, which is a card
— background, border, shadow — and needs inner padding. The hero search
is not a card: .heroMode strips all of it, padding included.

Both selectors are specificity (0,1,0), so source order decides, and
.heroMode only wins because it is declared right after .filterBar. A
bare .filterBar rule inside a media query comes later and silently wins
instead. The two rules directly below this one in the same block were
already written as `.filterBar:not(.heroMode)`; this one was missed.

Scoping it aligns the search box, hint and location link to the same
left edge as the headline and gives the field back its 28px.

The location link also carried its own 6px of button padding, so its
label started further right than the hint even once the boxes agreed.
Pulled back with a negative margin, which keeps the tap target.

The guard is a stylesheet test: the failure is a plausible-looking
layout rather than a broken one, so nothing short of measuring or
looking would catch it. Verified by reverting: it names ".filterBar sets
padding".

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-26 21:32:41 +01:00
tudor d55f6cce23 Merge pull request 'fix(suggest): let the dropdown out of the hero panel' (#128) from fix/hero-dropdown-clipping into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 50s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 5m50s
Reviewed-on: #128
2026-08-26 20:23:33 +00:00
TudorandClaude Opus 5 3236efa846 fix(map): the hero map's fade to the header was hardcoded white
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m7s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 10s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 11s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 49s
The fade between the map band and the school header ramped through
rgba(255,255,255,...) and landed on var(--bg-card). In the light theme
that is white into white and invisible, as designed. In the dark theme
it climbed to 95% WHITE and then met a near-black card, putting a bright
band across the full width exactly where the map should dissolve into
the title.

Fading to the colour the gradient lands on is the whole trick, and it
only works if that colour is a token — so --bg-card-rgb now exists in
both theme blocks, matching the --hero-ground-rgb precedent.

Two more defects in the same file, same cause, found while in there:

The controls floating over the map paired a hardcoded white background
with color: var(--text-primary), which resolves to #E9EEF0 in dark —
near-white text on a near-white button. These deliberately do NOT follow
the theme, because the map tiles are light in both, so the ink is now
literal too and says why. A themed token is the wrong tool for a surface
that never changes.

The loading skeleton swept 50% white across var(--bg-secondary), which
is a bright flash every 1.4s on a dark page. It now sweeps toward the
card colour, a shade lighter than the ground in both themes.

The guard is a stylesheet test rather than a render test, because the
bug is invisible in the theme it was written for. Verified by reverting
each fix in turn: it names .fade and .openHint exactly.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-26 21:20:57 +01:00
TudorandClaude Opus 5 d5a6db289d fix(suggest): let the dropdown out of the hero panel
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 11s
PR Checks / Build Frontend (no push) (pull_request) Successful in 45s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 54s
.heroPanel had overflow: hidden to clip its artwork and scrim to the
rounded corners. It clipped the suggestion dropdown too. Measured on
staging with the flag on: the list runs 482 to 802, the panel ends at
624 — so 178px of 320 was cut off, about half the options, with nothing
on screen to say anything was missing.

The two things that actually needed clipping now round themselves:
.heroArt gets border-radius: inherit plus its own overflow, and the
::before scrim inherits the radius. Below 860px the artwork is a band
flush with the top of the panel rather than a layer covering it, so it
takes the top two corners only — inheriting all four would leave it
floating with rounded corners against the copy.

Nothing else depended on the panel clipping: .valueProps below it is
entirely static, so a positioned dropdown paints above it without a
z-index fight.

The regression test asserts the LAST option is the element actually
painted at its own coordinates. toBeVisible() would not have caught
this — it checks for a non-empty box and visibility, and an ancestor's
overflow clips neither. elementFromPoint catches clipping and occlusion
alike.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-26 21:10:32 +01:00
tudor d8ccb5b733 Merge pull request 'feat(suggest): school autosuggest, and the rate-limit fix it needed first' (#127) from feat/school-autosuggest into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 20s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 51s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 5m49s
Reviewed-on: #127
2026-08-26 19:58:57 +00:00
TudorandClaude Opus 5 0fa1a292c7 fix(api): bound what a forged CF-Connecting-IP can buy
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 18s
PR Checks / Build Frontend (no push) (pull_request) Successful in 45s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 10s
Code review, both findings valid.

The design doc claimed Cloudflare "replaces the header, so a browser
cannot forge it", and that only the X-Forwarded-For fallback was
forgeable. That is true only for traffic that actually passed through
Cloudflare, and nothing in this process can verify that it did. Reaching
the origin directly, both headers are equally attacker-controlled — and
rotating CF-Connecting-IP mints a fresh rate-limit bucket per request,
defeating per-client limits on every endpoint including the
DataFrame-heavy /api/schools. Against abuse that is worse than the
shared bucket it replaced, which at least capped everyone together.

So the ceiling comes back. I dropped it earlier arguing it belonged at
Cloudflare; that argument assumed the keying was sound, and it is not.
GlobalRateLimitMiddleware counts all /api/ traffic in a fixed window
against a total, independent of client identity, outermost so it refuses
before any work happens. Written by hand because slowapi cannot express
a global cap: default_limits and application_limits are both keyed by
key_func, and the latter needs middleware this app does not install.

It does not make the header trustworthy — it makes trusting it
survivable. The real fix is Authenticated Origin Pulls or an origin
firewall, now documented in DEPLOY.md as the open gap it is.

127.0.0.1 is exempt: the healthcheck curls localhost from inside the
container, and starving it would restart the container and turn a load
spike into an outage loop. Keyed on the peer address, never the Host
header, which the caller sets.

Second finding: suggest_schools_typesense promised "never raises" while
the parsing loop sat outside the try, so int(None) on a malformed
document would have made a keystroke a 500. The loop now skips bad rows
rather than dropping the whole list — and a hit with no document no
longer becomes a suggestion pointing at /school/0.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-26 20:56:32 +01:00
tudor c3f044bd65 Merge pull request 'docs(flags): Unleash does not create flags by itself' (#126) from fix/flags-runbook into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 44s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 53s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 2m10s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 1m46s
Reviewed-on: #126
2026-08-26 19:48:58 +00:00
TudorandClaude Opus 5 d2115364ae test(e2e): autosuggest journeys, gated on the flag
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 31s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 1m13s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 2m40s
Feature state is read from its observable effect — whether the search
box is a combobox — because /api/flags is denied to the public on
purpose. Same approach as the distance journeys.

The three endpoint tests are ungated: /api/suggest is live whether or
not the UI is, which is what lets it be smoke-tested in an environment
where the feature is still dark.

The flag-off journey asserts the plain search still works, not just that
the combobox is absent. Verified against staging, where the flag is off:
it passes and the flag-on journey correctly skips.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-26 20:39:33 +01:00
TudorandClaude Opus 5 28cf0a342c feat(suggest): wire autosuggest into the search box behind a flag
Off means off — no combobox role, no listener, no fetch. A test asserts
the absence of the request, not just the absence of the dropdown,
because a hidden-but-fetching control would still be spending the rate
limit on a feature nobody can see.

Enter with no active option falls through to the form's submit handler
and searches the typed text exactly as before. The existing behaviour is
preserved, not replaced, and that has its own test.

Suppressed once the value parses as a postcode: the box takes a name OR
a postcode, and suggesting schools during postcode entry fights the user.

.omniBoxContainer gains position: relative — the dropdown is absolutely
positioned and without it would have anchored to the page instead.

Four render sites, all wired: page.tsx renders HomeView in the success
path AND the catch fallback, and HomeView renders FilterBar as hero AND
sticky. Missing any one would make the flag silently do nothing
somewhere.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-26 20:38:52 +01:00
TudorandClaude Opus 5 d88e77f459 feat(suggest): the dropdown, with combobox ARIA
Presentational only — it fetches nothing and owns no state, so the
fetching rules and the ARIA rules can be read separately.

onMouseDown, not onClick. The input's blur handler closes the list and
blur fires before click, so a click handler never runs: the classic bug
where a dropdown works perfectly by keyboard and is dead to the mouse.

The plan's CSS guessed at token names like --color-surface. The real
tokens are --bg-card, --border, --text-muted, --bg-secondary and
--shadow-soft, and all five are redefined in the dark theme — invented
names would have silently fallen back to hardcoded light values and
broken dark mode.

Local authority is rendered because there are many schools called
'St Mary's'; a list without it is unusable for exactly the query
autosuggest exists to serve, which is what the test asserts.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-26 20:36:39 +01:00
TudorandClaude Opus 5 06eb433db5 feat(suggest): debounced, abortable suggestion hook
The AbortController is correctness, not economy. Without it a slow
response for 'st' can land after the fast one for 'st marys' and replace
a correct list with a stale one — the classic autosuggest race.

No cache: 'no-store', unlike the compare modal's search. This is the one
endpoint where prefix queries repeat most across users, so discarding
the browser cache and the backend's ETag 304s would be throwing away the
cheapest win available.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-26 20:35:27 +01:00
TudorandClaude Opus 5 1a6d349dad feat(suggest): GET /api/suggest, cacheable and DataFrame-free
A dedicated endpoint rather than a mode of /api/schools, because that
path filters and sorts 25,000 pandas rows per query while holding the
GIL — affordable once per search, not once per keystroke. A test asserts
the distinction directly by making load_school_data raise and requiring
the endpoint to answer anyway.

Nothing errors on ordinary input: a short query, no matches, or
Typesense being down are all 200 with an empty list.

Cached deliberately. Prefix queries repeat enormously across users and
school names change once a year, so s-maxage plus the existing ETag
middleware turns most keystrokes into 304s.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-26 20:34:34 +01:00
TudorandClaude Opus 5 75d3534d82 feat(suggest): Typesense rows for autosuggest, no DataFrame
search_schools_typesense returns URNs, which forces the caller to
hydrate from the 25,000-row in-memory frame. Every field a suggestion
needs is already in the Typesense document, so this returns documents
and the caller needs no pandas at all — the difference between a query
that can run per keystroke and one that cannot.

Never raises. Typesense unreachable or erroring gives an empty list,
because a dropdown that quietly stops appearing is the right failure for
a keystroke path.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-26 20:33:42 +01:00
TudorandClaude Opus 5 ff041544f2 fix(api): rate-limit per caller, not per proxy
The limiter keyed on request.client.host, which in staging and prod is
the Next container — the backend has no published ports and nothing else
can reach it. So every browser user on the site shared one 60/minute
bucket per route. Measured against staging: 70 concurrent requests to
/api/schools returned exactly 60 OK and 10 refused, from one machine.

CF-Connecting-IP first. Cloudflare fronts both environments and
overwrites any client-supplied value, which a parsed X-Forwarded-For
chain does not guarantee. The XFF fallback is forgeable only from inside
the Docker network.

Named rather than hidden: the shared bucket was an accidental global
throttle on a single-process backend, and correct per-user keying
removes it. A real global ceiling belongs at Cloudflare, which is
already in the path; slowapi cannot express one without a second Limiter
and middleware this app does not install.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-26 20:32:58 +01:00
TudorandClaude Opus 5 59265f78b6 docs(flags): Unleash does not create flags by itself
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m6s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 35s
PR Checks / Build Frontend (no push) (pull_request) Successful in 46s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 1m15s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 35s
The runbook said a flag 'appears in the Unleash UI after the backend has
evaluated it once'. That is wrong. SDKs read definitions from the server
and never register anything, and metrics for an unknown flag are
discarded — so a declared flag is evaluated on every request, stays
False forever, and never shows up until someone creates it by hand.

Found the way these things usually are: staging had been running the
flag code for a while and the UI was still empty.

Also names the environment trap while here — each stack's token is
scoped to one environment, so toggling the other does nothing visible
and looks like the flag is broken.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-26 20:11:45 +01:00
TudorandClaude Opus 5 6e0a278340 docs(suggest): implementation plan, eight tasks
Self-review caught three defects in the plan. .omniBoxContainer, the
wrapper the dropdown positions against, does not declare position:
relative — without it the list anchors to the page. The postcode
suppression test typed character by character, so it would have asserted
no request while 'NW1' legitimately fires one; it now sets the value in
one go. And the Enter-submits-search assertion needed waitFor, because
updateURL pushes inside startTransition.

Task 1 is the one to review hardest: it is the only unflagged change and
it alters rate limiting for every endpoint.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-26 19:48:39 +01:00
TudorandClaude Opus 5 e651dd0d65 docs(suggest): the in-app global ceiling would not have worked
Reading slowapi rather than assuming: default_limits and
application_limits are both evaluated with the same key_func, so they
are per-client across routes, not global. And application_limits only
apply 'if in_middleware' — this app installs no SlowAPIMiddleware, so
they would never have fired at all.

A genuine global cap would need a second Limiter with a constant key
plus that middleware. Cloudflare is already in the path on both
environments and does this at the right layer, so the ceiling is named
as a follow-up there rather than built badly here.

The risk that leaves is stated plainly in the risks section instead of
being papered over with a mechanism that does not do the job.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-26 19:44:21 +01:00
TudorandClaude Opus 5 22c113fc29 docs(suggest): design for school autosuggest
The load-bearing finding is not about autosuggest. The rate limiter keys
on request.client.host, which in staging and prod is the Next container
— so all browser users share one 60/min bucket per route. Measured
against staging: 70 concurrent requests gave exactly 60 x 200 and
10 x 429. Eight concurrent searchers would 429 the site once each
keystroke costs a request, so the keying fix is part of this work.

Both environments are behind Cloudflare, which sets CF-Connecting-IP and
overwrites any client-supplied value — trustworthy in a way a parsed
X-Forwarded-For chain is not, and the backend is unreachable except
through the Next proxy.

Named honestly: the shared bucket has been an accidental global throttle
on a single-process backend, so correct per-user keying removes a
protection. A global ceiling ships with it rather than instead of it.

Suggestions come from Typesense alone. The existing search path filters
a 25,000-row DataFrame per query, which is exactly the cost a keystroke
endpoint cannot pay, so there is deliberately no DataFrame fallback.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-26 19:41:00 +01:00
tudor e953ee7c5f Merge pull request 'feat(flags): ship-dark feature flags, with last-distance-offered behind the first one' (#125) from feat/feature-flags into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 40s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 50s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 1m34s
Reviewed-on: #125
2026-08-23 11:34:56 +00:00
TudorandClaude Opus 5 413d86cc3c chore(flags): wire UNLEASH_URL and the SDK cache volume into the stacks
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 10s
PR Checks / Build Backend (no push) (pull_request) Successful in 27s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 2m59s
Both variables default to empty, so an environment without Unleash has
every flag off — the correct dark state rather than a boot failure.

The cache volume is the mitigation for the one real regression risk in
this design: the SDK evaluates everything False until it syncs, so a
backend cold-starting with an empty cache while Unleash is unreachable
would make a *released* feature disappear. On a named volume the disk
cache survives a restart.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-23 10:58:29 +01:00
TudorandClaude Opus 5 4f01fbdedb test(e2e): make the distance journeys fail loudly, not skip quietly
The existing distance journeys all skip when no school has a published
figure, which is right when the feature is off — and wrong when it is
supposed to be on and is silently broken, because that shows up as a
green run full of skips. The new gate fails in exactly that case.

Feature state is read from the data, not from /api/flags: the public
proxy denies that path on purpose, since it names unreleased features.
Presence of the admission_distance key is the observable effect.

Verified against staging, where the feature is currently on: the on-gate
passes, the off-gate skips, the existing eight distance journeys are
unaffected.

One honest caveat — the /api/flags check passes on staging today because
that image predates the endpoint, not because the denylist works. The
denylist itself is covered by the jest unit test; this is defence in
depth and becomes a real assertion once deployed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-23 10:57:34 +01:00
TudorandClaude Opus 5 c3ba7aae0d feat(flags): ship last-distance-offered dark behind a flag
One gate, at the source. The frontend needs no change: DistanceSection
already returns null when distance_m is missing, and the admissions
block already conditions on (admissions || admissionDistance). Only 57
local authorities publish cut-offs, so the off-path is the commonest
path on the site and is well covered already.

Absent, not null. /api/schools/ is public and unauthenticated, so a
field left in the payload is a published field — the reasoning already
recorded in c9a1892 when history was withheld. The two are also
different claims: null says this school has no cut-off, absent says
cut-offs are not being published at all. The frontend type now says so.

The feature is on main and live on staging and has never reached
production, which is what makes it the right first consumer: the flag
lets the code promote without the feature appearing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-23 10:56:33 +01:00
TudorandClaude Opus 5 54a30de0d8 feat(flags): server-side getFlags for the frontend
Ships without a consumer, deliberately. The first flag needs none — the
backend withholds the field and the page follows — but 'UI elements on
existing pages' is one of the three surfaces this capability exists for,
and a flag layer that cannot gate one is incomplete.

Never throws: an unreadable flag is a dark one, which matches the
backend's fail-closed default. A page that 500s because the flags
endpoint blinked would be a worse outcome than a hidden feature.

Reading flags pins the calling route to a 300s ISR floor, since Next
takes the lowest revalidate among a route's fetches. That matches what
/school/[slug] already sits at, and it is the same property that makes a
flip propagate without a webhook.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-23 10:55:32 +01:00
TudorandClaude Opus 5 c30ad1db07 feat(flags): serve /api/flags, and keep the public proxy off it
The endpoint and its exposure control ship together on purpose. The
moment /api/flags exists, app/api/[...path] forwards it — and the
response names every unreleased feature the codebase knows about, along
with whether it is on. Publishing that is the opposite of shipping dark.

Denied on an exact first-segment match, not a prefix, so /api/flagship
does not go down with /api/flags. Next reads the endpoint server-side
over the Docker network, which never transits the public proxy.

jest.setup.js now guards its browser globals. It runs for every suite,
including the one that declares @jest-environment node to exercise the
route handler — NextRequest needs Fetch API globals jsdom lacks, and
there is no window there to define matchMedia on.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-23 10:54:26 +01:00
TudorandClaude Opus 5 7424cef7c6 feat(flags): the registry and a fail-closed Unleash client
Unleash holds flag state; it does not hold the list of flags. REGISTRY
is that list, because the SDK evaluates an unknown flag to False and
without a registry that is an undeclared False — indistinguishable from
a typo in a flag name.

Fail-closed throughout, and never raises: an unset UNLEASH_URL, an
unreachable server, a client that throws, an undeclared name — all
False. A flag layer that can 500 a request path or stop the API booting
is worse than one that is switched off.

Every flag defaults to False, with no per-flag override, because a flag
that defaults on is a kill switch and this is deliberately not one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-23 10:52:55 +01:00
TudorandClaude Opus 5 01ccbb8e82 feat(flags): add the Unleash stack and its runbook
Its own Portainer stack, belonging to neither application stack: a
staging redeploy must not be able to disturb production's flag state.

One instance serves both. OSS Unleash ships development and production
environments with environment-scoped client tokens, so the same flag
holds independent state in each — which is what lets a feature be on in
staging, where the E2E journeys exercise it, while production stays dark.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-23 10:51:35 +01:00
TudorandClaude Opus 5 c339c2f1a1 docs(flags): implementation plan, eight tasks
Task 1 is the Unleash stack and ends with a human step — the Portainer
deploy and the token generation cannot be automated from here. Nothing
else blocks on it: an unset UNLEASH_URL means every flag is False, which
is what local development and CI get, so the whole suite runs without a
flag server existing.

Self-review caught three defects in the plan itself. get_supplementary_data
takes (db, urn), not (urn), and the test DataFrame was minimised to the
point where the endpoint would have failed for reasons unrelated to
flags — both now copy the known-good shape from test_school_details.py.
The proxy test needs the node jest environment, since NextRequest wants
Fetch API globals jsdom does not provide. And the e2e off-state check
hardcoded a URN, so a 404 page would have satisfied it without proving
anything.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-23 10:46:15 +01:00
TudorandClaude Opus 5 e2ca3d79f9 docs(flags): drop the webhook — the seven-day premise was wrong
Next uses the LOWEST revalidate among a route's fetches, not the segment
value. School pages fetch school details at 300s and place pages fetch
national averages at 3600s, so the effective ISR period is five minutes
and one hour respectively — not the seven days the segment declares.

A flag flip therefore propagates on its own, well inside the monthly,
by-hand cadence these flags are for. That deletes two webhook
integrations, a revalidate route, a secret-in-query-string scheme, an
idempotency requirement, and the rule that every fetch carry a cache
tag — which was the part most likely to rot as fetches are added.

Two constraints survive: a flag must never gate content on a
force-static page, because app/admissions never revalidates; and a
route-family flag must rebuild the sitemap, deferred with the route case
since no flag in scope touches it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-23 10:40:35 +01:00
TudorandClaude Opus 5 c2364bf09e docs(flags): design for a ship-dark feature flag layer
Unleash self-hosted in its own Portainer stack, with FastAPI holding the
only SDK and Next reading flags through a tagged fetch.

The two hard parts are consequences of putting flag state in a service
rather than the repo: main stops being the whole truth about what is on,
and a flag can now change without the deploy that would have cleared the
caches. A code-declared registry bounds the first; webhook-driven
revalidateTag handles the second.

Cache tagging is deliberately coarse — every server fetch carries the
flags tag, not just the flags fetch itself. The first consumer proves
why: admission_distance changes the shape of /api/schools/{urn}, so a
narrow purge would leave ~25,000 school pages serving the pre-flip
render for a week, invisibly.

First consumer is the last-distance-offered feature, which is on main
and staging and has never reached production. It needs one gate, at the
API, because the frontend already no-ops on a missing field.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-23 09:47:58 +01:00
tudor 43e0621728 Merge pull request 'fix(places): phase links must stay in their own namespace' (#124) from fix/place-phase-links into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 19s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 50s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 1m31s
Reviewed-on: #124
2026-08-22 17:33:12 +00:00
TudorandClaude Opus 5 d1358cc00f fix(places): phase links must stay in their own namespace
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 17s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 2m37s
Every place page built its phase links as /schools/[slug]/[phase], the
shape that belongs to towns alone.

On an authority page that pointed into the town namespace. For 87 of the
151 authorities the target does not exist and the link 404s; for the
other 64 it resolves to the town of the same name — a different set of
schools, which is precisely the near-duplicate the two namespaces were
introduced to prevent. On an outcode page it 404s outright.

Two causes behind it, both a rule written twice and inherited by only
one of the places that needed it.

The authority phase route was in the spec and dropped by the plan, which
built the three bare routes and no fourth. The sitemap is generated from
the place registry, which was right about them all along, so 302
authority phase URLs have been submitted to Google and every one 404s.
Adding the route makes the sitemap true and serves a real query —
admissions are authority-run, so "primary schools in Kent" is how a
parent searches before they have settled on a town.

The outcode variants were the opposite: the registry computed phases for
outcodes although the spec gives them no route, and the sitemap knew to
skip them while the API did not. The registry now decides alone, and the
sitemap's duplicate of that rule is gone.

Also: an authority under the five-school threshold has no page, so the
API sends a null slug for it and the page names it without linking.
Two English authorities are in that position. It was unreachable in
today's data — verified across the EC and TR outcodes — but the thin
place redirect would have sent a reader to a 404 the year it isn't.

The e2e journey now walks every /schools link a page of each family
emits and requires a 200, which is the check that was missing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-22 17:22:11 +01:00
tudor 865a69b54d Merge pull request 'feat(places): list schools alphabetically on place pages' (#123) from feat/place-alphabetical-sort into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 20s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 50s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 1m27s
Reviewed-on: #123
2026-08-21 23:22:49 +00:00
TudorandClaude Opus 5 9cc87c41bb fix(places): a phase page needs results, not merely publishable schools
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m2s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 16s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 2m4s
Asked where schools with no results should sit in an alphabetical list, and
found that some pages were almost entirely made of them.

The per-phase threshold counted schools that were publishable — a result OR
an Ofsted grade — while a phase page exists for its results column.
/schools/kent/primary published with none of its five rows carrying a result;
Minehead had one of seven, Buntingford one of five. Forty-four phase pages
were majority-blank.

It is the same rule as "no page without a local average", which was written
into the spec as a thin-page control and never extended per phase.

The threshold now counts schools with a result for that phase. It gates
whether the page exists; it does not filter rows — a page that publishes still
lists every school of the phase, because someone looking up a school by name
has to find it whether or not it published results.

126 of 1,012 variant pages stop publishing: 62 primary, 64 secondary. Every
one of them was a table with too little in it to be worth a page.

The ordering itself is unchanged: pure A-Z, blanks interleaved. A school sits
where its name says it does, and at roughly a tenth of rows that reads fine.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-22 00:14:31 +01:00
TudorandClaude Opus 5 8967966eef feat(places): list schools alphabetically on place pages
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m2s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 16s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 11s
PR Checks / AI Code Review (Claude) (pull_request) Canceled after 1m21s
Someone on a place page is usually looking for a school they can name, so the
order should serve scanning for it rather than ranking. /api/rankings keeps
its league-table ordering; this is a place-page decision, not a site-wide one.
Sorted case-insensitively, or a capitalised name would sort ahead of every
lowercase one.

The change made five pieces of copy untrue, so they go with it. The phase
variant titled itself "— Ranked", and all four route families described
themselves as "ranked by SATs and GCSE results". A page that opens by claiming
an order it does not keep is worse than one that claims nothing.

The ItemList markup carried `position` with no declared order, which reads as
a ranking. It now declares ItemListOrderAscending, so the structured data says
what the table does.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-22 00:09:58 +01:00
tudor 4a9a5c734b Merge pull request 'fix(e2e): three assertions that were wrong about correct behaviour' (#122) from fix/e2e-canonical-and-robots into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 50s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Successful in 1m26s
Reviewed-on: #122
2026-08-21 23:07:57 +00:00
TudorandClaude Opus 5 4e82e6c916 fix(e2e): three assertions that were wrong about correct behaviour
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m2s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 11s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 41s
The staging gate was red on three journeys. All three were faults in the
tests; the site was behaving correctly in each case.

Next normalises canonical URLs against trailingSlash:false, so the homepage
ships "https://www.schoolcompare.co.uk" with no slash while every other route
keeps its path. Both address the same document. The test hardcoded the slash
and so failed only on the root — /rankings and /admissions passed throughout,
which is what made it look like a homepage bug rather than a test bug.
Compared with trailing slashes stripped from both sides.

The robots.txt assertion matched "Disallow: /" anywhere in the file and
tripped over the AI-crawler groups Cloudflare injects — ClaudeBot, GPTBot,
Amazonbot and six others all carry a blanket disallow, deliberately, and none
of them is Googlebot. It now parses the file into user-agent groups and checks
only the "*" group, which is also the thing the test was always trying to say:
Google may crawl the page, so it can see the noindex header.

Both were the same mistake as the doubled brand: asserting a naive string
rather than the semantics, and asserting against what the code assembles
rather than what the page renders.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 23:51:34 +01:00
tudor d4340a8fdd Merge pull request 'feat(places): name every authority a place sits in' (#121) from feat/place-multiple-authorities into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 19s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 50s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 1m30s
Reviewed-on: #121
2026-08-21 21:56:50 +00:00
TudorandClaude Opus 5 bb2f7a5841 fix(places): address review, and merge places GIAS spells more than one way
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m2s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 17s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 2m34s
Two findings from review on #121, plus a third the review prompted.

The cap at three authorities silently dropped the fourth in exactly the case
where the information matters most — a genuinely fragmented place — and
contradicted the stated goal of naming every authority a place sits in. It is
gone. The share rule was always the real limit and already bounds the list at
ten. Measured against the live corpus, one town would have been truncated
today: LONDON, split evenly between Hackney, Lambeth, Westminster and
Lewisham.

parent_authority used mode() while authorities used value_counts(), and on an
exact tie pandas does not guarantee the two pick the same name, so the 301
could have pointed somewhere other than the authority named first on the page.
The parent is now derived from authorities[0]: one computation, one answer.
It also inherits the sentinel filter, so a place can no longer redirect to
/schools/authority/does-not-apply.

Chasing the truncation case surfaced a worse bug. Places were grouped by raw
town value, but the registry is keyed by slug, and GIAS spells the same place
several ways. Five town slugs come from more than one spelling: "London"
(1,819 schools) and "LONDON" (12) both slugify to `london`, so the later group
simply overwrote the earlier one — /schools/london could have shown twelve
schools, silently, depending on row order. Weston-super-Mare was split 14/19
across two spellings and Newcastle-under-Lyme across three. Grouping is now by
slug, and the display name is the most common spelling.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 22:42:18 +01:00
TudorandClaude Opus 5 1cb5314c53 feat(places): name every authority a place sits in
SW19 is mostly Merton but partly Wandsworth, and the page said only Merton.
The cause was one field doing two jobs: _parent_authority takes the modal
authority, which is right for a 301 target and wrong as a statement about
where a place is.

This is not a corner case. A quarter of viable outcodes (425 of 1,760) and a
third of viable towns (263 of 783) cross an authority boundary — Bedford the
town spans Bedford and Central Bedfordshire.

Place now carries `authorities`, every authority holding at least a tenth of
the schools and at least two of them, largest first. parent_authority stays
single and unchanged, because a redirect still needs one target.

The share threshold exists because GIAS carries postcode errors: EN6 lists two
Shropshire schools among fourteen in Hertfordshire, and a bare "any authority
present" rule would print those as though they were real. A place too small or
too fragmented to clear the threshold still names its largest, so the page
never goes silent about where it is.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 22:40:58 +01:00
tudor 4cea26b813 Merge pull request 'fix(places): align the measure column's heading with its values' (#120) from fix/place-table-alignment into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 49s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 1m30s
Reviewed-on: #120
2026-08-21 21:21:03 +00:00
171 changed files with 25734 additions and 361 deletions

No files matched your search

+203 -15
View File
@@ -6,6 +6,7 @@ Uses real data from UK Government Compare School Performance downloads.
import hashlib import hashlib
import re import re
import time
from contextlib import asynccontextmanager from contextlib import asynccontextmanager
from datetime import datetime, timezone from datetime import datetime, timezone
from typing import Optional from typing import Optional
@@ -15,7 +16,7 @@ import pandas as pd
from fastapi import FastAPI, HTTPException, Query, Request, Depends, Header from fastapi import FastAPI, HTTPException, Query, Request, Depends, Header
from fastapi.middleware.cors import CORSMiddleware from fastapi.middleware.cors import CORSMiddleware
from fastapi.middleware.gzip import GZipMiddleware from fastapi.middleware.gzip import GZipMiddleware
from fastapi.responses import FileResponse, Response from fastapi.responses import FileResponse, JSONResponse, Response
from fastapi.staticfiles import StaticFiles from fastapi.staticfiles import StaticFiles
from slowapi import Limiter, _rate_limit_exceeded_handler from slowapi import Limiter, _rate_limit_exceeded_handler
from slowapi.util import get_remote_address from slowapi.util import get_remote_address
@@ -33,8 +34,10 @@ from .data_loader import (
get_supplementary_data, get_supplementary_data,
get_supplementary_data_batch, get_supplementary_data_batch,
search_schools_typesense, search_schools_typesense,
suggest_schools_typesense,
) )
from .data_loader import get_data_info as get_db_info from .data_loader import get_data_info as get_db_info
from . import flags
from .places import build_place_registry from .places import build_place_registry
from .schemas import METRIC_DEFINITIONS, RANKING_COLUMNS, SCHOOL_COLUMNS from .schemas import METRIC_DEFINITIONS, RANKING_COLUMNS, SCHOOL_COLUMNS
from .utils import clean_for_json, convert_to_native from .utils import clean_for_json, convert_to_native
@@ -222,10 +225,10 @@ def _place_sitemap_rows(kinds: tuple[str, ...]) -> list[str]:
if p.kind not in kinds: if p.kind not in kinds:
continue continue
rows.append(_url_element(BASE_URL + _place_url(p))) rows.append(_url_element(BASE_URL + _place_url(p)))
# Outcodes carry no phase variants: nobody searches "primary schools # Which phases a place publishes is the registry's decision alone —
# in SW11", so the routes do not exist to submit. # outcodes report none, because the spec gives them no phase route.
if p.kind == "outcode": # Repeating that rule here was how the page and the sitemap came to
continue # disagree about which URLs exist.
for phase in ("primary", "secondary"): for phase in ("primary", "secondary"):
if p.publishes_phase(phase): if p.publishes_phase(phase):
rows.append(_url_element(f"{BASE_URL}{_place_url(p)}/{phase}")) rows.append(_url_element(f"{BASE_URL}{_place_url(p)}/{phase}"))
@@ -295,8 +298,101 @@ def clean_filter_values(series: pd.Series) -> list[str]:
# SECURITY MIDDLEWARE & HELPERS # SECURITY MIDDLEWARE & HELPERS
# ============================================================================= # =============================================================================
# Rate limiter def client_key(request: Request) -> str:
limiter = Limiter(key_func=get_remote_address) """The rate-limit bucket: the real caller, not the proxy in front of them.
`get_remote_address` reads request.client.host. In staging and production
the backend has no published ports and sits on the internal network, so its
only caller is the Next proxy — meaning every browser user on the site
shared one bucket. Measured before this fix: 70 concurrent requests to
/api/schools returned 60 OK and 10 refused.
CF-Connecting-IP first, because Cloudflare (in front of both environments)
sets it on every origin request and *overwrites* any client-supplied value,
which a parsed X-Forwarded-For chain does not guarantee. The XFF fallback is
forgeable, but only by a caller already inside the Docker network, which is
the one place nothing untrusted can reach.
"""
cf = request.headers.get("cf-connecting-ip")
if cf:
return cf.strip()
xff = request.headers.get("x-forwarded-for")
if xff:
return xff.split(",")[0].strip()
return get_remote_address(request)
# Per-client limiter. Paired with the global ceiling below — the two do
# different jobs and neither substitutes for the other.
limiter = Limiter(key_func=client_key)
# --- The ceiling no header can raise ----------------------------------------
#
# client_key trusts CF-Connecting-IP, and nothing in this process can tell an
# edge-set header from an attacker-set one. That distinction can only be made
# at Cloudflare, with Authenticated Origin Pulls or an origin firewall. A
# caller reaching the origin directly could otherwise mint a fresh rate-limit
# bucket per request and evade per-client limits entirely — which would make
# correct keying a net regression against abuse, since the single shared bucket
# it replaced at least capped everyone at 60/minute together.
#
# So per-client limits give fairness, and this gives the origin a hard total.
# It does not make the header trustworthy; it bounds what trusting it can cost.
# The header problem itself is closed at Cloudflare, not here.
#
# [window_start_monotonic, count], or None before the first request. A fixed
# window is crude, which is right for a backstop: it has to be obviously
# correct rather than fair.
_global_window: Optional[list] = None
# The container healthcheck runs `curl http://localhost:80/api/data-info` from
# inside the container. Starving it would fail the check, restart the
# container, and turn a load spike into an outage loop — the ceiling exists to
# protect the origin, not to kill it.
_LOCAL_HOSTS = frozenset({"127.0.0.1", "::1", "localhost"})
def exempt_from_ceiling(request: Request) -> bool:
"""Whether the ceiling should ignore this request.
Its own function so the rule is testable without standing up a server —
and so the healthcheck exemption is somewhere a reader can find it.
"""
if not request.url.path.startswith("/api/"):
return True
# The peer address, never the Host header: Host is set by the caller and
# would hand every attacker an exemption.
return (request.client.host if request.client else "") in _LOCAL_HOSTS
class GlobalRateLimitMiddleware(BaseHTTPMiddleware):
"""A cap on total /api/ traffic, independent of any client identity."""
async def dispatch(self, request: Request, call_next):
global _global_window
if exempt_from_ceiling(request):
return await call_next(request)
now = time.monotonic()
# One event loop, and no await between the read and the write, so this
# sequence is atomic without a lock.
if _global_window is None or now - _global_window[0] >= 60:
_global_window = [now, 0]
_global_window[1] += 1
if _global_window[1] > settings.global_rate_limit_per_minute:
return JSONResponse(
# Distinguishable from slowapi's per-client 429: an operator
# reading logs has to be able to tell "one noisy client" from
# "the origin is saturated".
{"detail": "The service is at capacity. Please retry shortly."},
status_code=429,
headers={"Retry-After":
str(max(1, int(60 - (now - _global_window[0]))))},
)
return await call_next(request)
class SecurityHeadersMiddleware(BaseHTTPMiddleware): class SecurityHeadersMiddleware(BaseHTTPMiddleware):
@@ -354,6 +450,7 @@ CACHE_RULES: list[tuple[str, tuple[int, int, int]]] = [
("/api/schools/", (300, 3600, 86400)), # /api/schools/{urn} ("/api/schools/", (300, 3600, 86400)), # /api/schools/{urn}
("/api/rankings", (60, 600, 3600)), ("/api/rankings", (60, 600, 3600)),
("/api/compare", (60, 600, 3600)), ("/api/compare", (60, 600, 3600)),
("/api/suggest", (60, 3600, 86400)), # autosuggest
("/api/schools", (30, 300, 1800)), # search list ("/api/schools", (30, 300, 1800)), # search list
] ]
@@ -459,6 +556,7 @@ def validate_postcode(postcode: Optional[str]) -> Optional[str]:
async def lifespan(app: FastAPI): async def lifespan(app: FastAPI):
"""Application lifespan - startup and shutdown events.""" """Application lifespan - startup and shutdown events."""
global _sitemaps global _sitemaps
flags.init()
print("Loading school data from marts...") print("Loading school data from marts...")
df = load_school_data() df = load_school_data()
if df.empty: if df.empty:
@@ -500,6 +598,10 @@ app.add_middleware(CacheAndETagMiddleware)
app.add_middleware(SecurityHeadersMiddleware) app.add_middleware(SecurityHeadersMiddleware)
app.add_middleware(RequestSizeLimitMiddleware) app.add_middleware(RequestSizeLimitMiddleware)
app.add_middleware(GZipMiddleware, minimum_size=512) app.add_middleware(GZipMiddleware, minimum_size=512)
# Added last, so it is outermost and refuses before anything downstream does
# work. A ceiling that only applies after the expensive part has run is not a
# ceiling.
app.add_middleware(GlobalRateLimitMiddleware)
# CORS middleware - restricted for production # CORS middleware - restricted for production
app.add_middleware( app.add_middleware(
@@ -806,11 +908,18 @@ async def get_school_details(request: Request, urn: int):
"census": supplementary.get("census"), "census": supplementary.get("census"),
"admissions": supplementary.get("admissions"), "admissions": supplementary.get("admissions"),
"admissions_history": supplementary.get("admissions_history") or [], "admissions_history": supplementary.get("admissions_history") or [],
"admission_distance": supplementary.get("admission_distance"), # Behind a flag, and withheld at the source rather than rendered-but-
# hidden: this endpoint is public and unauthenticated, so a field left
# in the payload is a published field. The key is absent, not null —
# null would state that this school has no cut-off, which is a
# different claim from "we are not publishing cut-offs".
**({"admission_distance": supplementary.get("admission_distance")}
if flags.is_enabled("admission_distance") else {}),
"sen_detail": supplementary.get("sen_detail"), "sen_detail": supplementary.get("sen_detail"),
"phonics": supplementary.get("phonics"), "phonics": supplementary.get("phonics"),
"deprivation": supplementary.get("deprivation"), "deprivation": supplementary.get("deprivation"),
"finance": supplementary.get("finance"), "finance": supplementary.get("finance"),
"destinations": supplementary.get("destinations"),
} }
@@ -1200,7 +1309,8 @@ async def get_place(request: Request, kind: str, slug: str,
if kind not in VALID_PLACE_KINDS: if kind not in VALID_PLACE_KINDS:
raise HTTPException(status_code=404, detail="No such place") raise HTTPException(status_code=404, detail="No such place")
place = get_place_registry().get(f"{kind}:{slug}") registry = get_place_registry()
place = registry.get(f"{kind}:{slug}")
if place is None: if place is None:
raise HTTPException(status_code=404, detail="No such place") raise HTTPException(status_code=404, detail="No such place")
@@ -1212,10 +1322,15 @@ async def get_place(request: Request, kind: str, slug: str,
if wanted and "phase" in rows.columns: if wanted and "phase" in rows.columns:
rows = rows[rows["phase"].fillna("").str.lower().isin(wanted)] rows = rows[rows["phase"].fillna("").str.lower().isin(wanted)]
# The metric the page ranks on, which is also the one it averages. # The metric the page shows, and averages.
metric = "attainment_8_score" if phase == "secondary" else "rwm_expected_pct" metric = "attainment_8_score" if phase == "secondary" else "rwm_expected_pct"
if metric in rows.columns:
rows = rows.sort_values(metric, ascending=False, na_position="last") # Alphabetical, not by score. A place page is read by someone looking for
# a school they can name, and scanning for it is what the order should
# serve. /rankings is where the league-table ordering lives, and it keeps
# sorting by metric.
if "school_name" in rows.columns:
rows = rows.sort_values("school_name", key=lambda c: c.str.lower())
averages = { averages = {
m: (None if m not in rows.columns or rows[m].dropna().empty m: (None if m not in rows.columns or rows[m].dropna().empty
@@ -1223,15 +1338,44 @@ async def get_place(request: Request, kind: str, slug: str,
for m in ("rwm_expected_pct", "attainment_8_score") for m in ("rwm_expected_pct", "attainment_8_score")
} }
cols = [c for c in SCHOOL_COLUMNS + ["latitude", "longitude", "phase", # dict.fromkeys, not a list: SCHOOL_COLUMNS already ends with latitude and
"rwm_expected_pct", "attainment_8_score", # longitude, so concatenating them again selected each twice and pandas
"total_pupils"] # dropped one of every duplicated pair with a "columns are not unique"
# warning. Ordered de-duplication keeps the column order and the warning
# cannot come back.
#
# nursery_provision and parliamentary_constituency are not in
# SCHOOL_COLUMNS and the place table shows both. The `in rows.columns`
# guard is what keeps a mart the pipeline has not rebuilt working: those
# two are the optional GIAS columns data_loader degrades to NULL.
cols = [c for c in dict.fromkeys(
SCHOOL_COLUMNS + ["latitude", "longitude", "phase",
"nursery_provision",
"parliamentary_constituency",
"rwm_expected_pct", "attainment_8_score",
"total_pupils"])
if c in rows.columns] if c in rows.columns]
return { return {
"place": {"kind": place.kind, "slug": place.slug, "name": place.name, "place": {"kind": place.kind, "slug": place.slug, "name": place.name,
"count": len(place.urns), "count": len(place.urns),
"parent_authority": place.parent_authority, "parent_authority": place.parent_authority,
# Every authority the place meaningfully sits in. SW19 is
# mostly Merton but partly Wandsworth; naming one asserts
# something false.
#
# The slug is null where that authority has no page of its
# own: City of London and the Isles of Scilly hold fewer
# schools than the threshold. Naming them is still right;
# linking them would be a 404.
"authorities": [
{"name": name,
"slug": (_slugify(name)
if f"authority:{_slugify(name)}" in registry
else None),
"count": n}
for name, n in place.authorities
],
# Only phases that clear the threshold, so the page links # Only phases that clear the threshold, so the page links
# variants that exist rather than 404s. # variants that exist rather than 404s.
"phases": [ph for ph in ("primary", "secondary") "phases": [ph for ph in ("primary", "secondary")
@@ -1241,6 +1385,50 @@ async def get_place(request: Request, kind: str, slug: str,
} }
# Two characters. One is not a query — it matches thousands of schools and the
# response is useless, so it is not worth a round trip.
SUGGEST_MIN_QUERY = 2
@app.get("/api/suggest")
@limiter.limit("120/minute")
async def suggest_schools(
request: Request,
q: str = Query("", max_length=100),
limit: int = Query(8, ge=1, le=20),
):
"""School name suggestions, from Typesense alone.
Deliberately not a mode of /api/schools: that path filters and sorts the
full in-memory DataFrame, which is far too expensive to run per keystroke.
Nothing here returns an error for ordinary input. A short query, no
matches, or Typesense being unreachable are all 200 with an empty list —
a dropdown that quietly does not appear is the right failure for a
keystroke path, and there is no DataFrame fallback because the 25,000-row
substring scan is precisely what this endpoint exists to avoid.
120/minute rather than the default 60: a 200 ms debounce makes typing
legitimately bursty.
"""
query = q.strip()
if len(query) < SUGGEST_MIN_QUERY:
return {"suggestions": []}
return {"suggestions": suggest_schools_typesense(query, limit)}
@app.get("/api/flags")
@limiter.limit(f"{settings.rate_limit_per_minute}/minute")
async def get_feature_flags(request: Request):
"""Every declared flag and its current value.
Internal only. The Next proxy denies this path, because the response names
every unreleased feature the codebase knows about — which is exactly what
shipping dark is meant to keep quiet.
"""
return flags.all_flags()
@app.get("/api/data-info") @app.get("/api/data-info")
@limiter.limit(f"{settings.rate_limit_per_minute}/minute") @limiter.limit(f"{settings.rate_limit_per_minute}/minute")
async def get_data_info(request: Request): async def get_data_info(request: Request):
+15
View File
@@ -35,6 +35,11 @@ class Settings(BaseSettings):
# Security # Security
admin_api_key: str = Field(default_factory=lambda: secrets.token_urlsafe(32)) admin_api_key: str = Field(default_factory=lambda: secrets.token_urlsafe(32))
rate_limit_per_minute: int = 60 # Requests per minute per IP rate_limit_per_minute: int = 60 # Requests per minute per IP
# A ceiling on total /api/ traffic, independent of any client identity.
# client_key trusts headers only Cloudflare can vouch for, so a caller
# reaching the origin directly could otherwise mint a fresh bucket per
# request. See GlobalRateLimitMiddleware in backend/app.py.
global_rate_limit_per_minute: int = 3000
rate_limit_burst: int = 10 # Allow burst of requests rate_limit_burst: int = 10 # Allow burst of requests
max_request_size: int = 1024 * 1024 # 1MB max request size max_request_size: int = 1024 * 1024 # 1MB max request size
@@ -42,6 +47,16 @@ class Settings(BaseSettings):
typesense_url: str = "http://localhost:8108" typesense_url: str = "http://localhost:8108"
typesense_api_key: str = "" typesense_api_key: str = ""
# Feature flags (Unleash). An empty unleash_url disables flags entirely and
# every flag evaluates False — the correct behaviour for local development
# and CI, and the reason no test needs a running Unleash.
unleash_url: str = ""
unleash_api_token: str = ""
unleash_app_name: str = "schoolcompare-backend"
# On a named volume, so a restart during an Unleash outage keeps
# last-known state instead of reverting a released feature to dark.
unleash_cache_directory: str = "/app/.unleash"
# Analytics # Analytics
ga_measurement_id: Optional[str] = "G-J0PCVT14NY" # Google Analytics 4 Measurement ID ga_measurement_id: Optional[str] = "G-J0PCVT14NY" # Google Analytics 4 Measurement ID
+298
View File
@@ -20,6 +20,7 @@ from .models import (
DimSchool, DimLocation, KS2Performance, DimSchool, DimLocation, KS2Performance,
FactOfstedInspection, FactAdmissions, FactAdmissionDistance, FactOfstedInspection, FactAdmissions, FactAdmissionDistance,
FactDeprivation, FactFinance, FactPupilCharacteristics, FactDeprivation, FactFinance, FactPupilCharacteristics,
FactKs4Destinations, FactKs5Destinations,
) )
from .ofsted_codes import ofsted_page_url, report_card_labels from .ofsted_codes import ofsted_page_url, report_card_labels
from .schemas import SCHOOL_TYPE_MAP from .schemas import SCHOOL_TYPE_MAP
@@ -100,6 +101,58 @@ def search_schools_typesense(query: str, limit: int = 250) -> List[int]:
return [] return []
# The most a public endpoint will return in one response.
SUGGEST_MAX_LIMIT = 20
# Fields a suggestion row carries, and the default when the document omits an
# optional one. phase and school_type are optional in the Typesense schema.
_SUGGEST_FIELDS = ("school_name", "local_authority", "postcode",
"phase", "school_type")
def suggest_schools_typesense(query: str, limit: int = 8) -> List[dict]:
"""Autosuggest rows straight from Typesense. Never raises.
Returns documents rather than URNs, unlike search_schools_typesense, so the
caller needs no DataFrame. Every field below is already in the index — see
pipeline/scripts/sync_typesense.py — which is what makes this cheap enough
to run per keystroke.
"""
client = _get_typesense_client()
if client is None:
return []
try:
result = client.collections["schools"].documents.search({
"q": query,
"query_by": "school_name,local_authority",
"per_page": max(1, min(limit, SUGGEST_MAX_LIMIT)),
"typo_tokens_threshold": 1,
})
except Exception:
# A dropdown that quietly stops appearing is the right failure here.
return []
rows = []
for hit in result.get("hits", []) or []:
doc = (hit or {}).get("document") or {}
try:
urn = int(doc["urn"])
except (KeyError, TypeError, ValueError):
# Skip the row, keep the rest. Typesense declares urn as int32 so
# this should be unreachable, but the index is a separate system
# that something other than this code can reindex — and "never
# raises" is a promise the keystroke path actually depends on.
# Dropping one malformed document is right; blanking the whole
# dropdown, or serving a suggestion pointing at /school/0, is not.
logging.getLogger(__name__).warning(
"skipping malformed suggestion document: %r", doc)
continue
row = {"urn": urn}
row.update({f: str(doc.get(f, "") or "") for f in _SUGGEST_FIELDS})
rows.append(row)
return rows
def normalize_school_type(school_type: Optional[str]) -> Optional[str]: def normalize_school_type(school_type: Optional[str]) -> Optional[str]:
"""Convert cryptic school type codes to user-friendly names.""" """Convert cryptic school type codes to user-friendly names."""
if not school_type: if not school_type:
@@ -764,6 +817,218 @@ def _finance_dict(f) -> dict:
} }
# Destination measures that are totals DfE published itself, rather than one of
# the categories that partition the cohort.
_AGGREGATE_MEASURES = {"agg_sustained_education", "agg_sustained_all"}
def _format_cohort_year(year) -> str | None:
"""202223 -> '2022/23'.
The section has to date its own cohort. Destination measures run about two
GCSE years behind the results shown above them on the same page, so an
undated figure reads as stale data rather than as a different question.
"""
if not year:
return None
text = str(year)
if len(text) == 6:
return f"{text[:4]}/{text[4:6]}"
if len(text) == 8:
return f"{text[:4]}/{text[6:8]}"
return text
_PUPIL_GROUPS = ("disadvantaged", "other", "all")
def _lone_hidden_groups(groups: dict) -> list:
"""Pupil groups hiding exactly one category — solvable by subtraction."""
return [
key for key, group in groups.items()
if sum(1 for c in group["categories"] if c["status"] == "suppressed") == 1
]
def _lone_hidden_categories(groups: dict) -> list:
"""Categories hidden in exactly one of several pupil groups."""
lone = []
categories = {c["category"] for g in groups.values() for c in g["categories"]}
for category in categories:
found = [
c for g in groups.values() for c in g["categories"]
if c["category"] == category
]
hidden = [c for c in found if c["status"] == "suppressed"]
if len(hidden) == 1 and len(found) > 1:
lone.append(category)
return lone
def disclosure_invariant_holds(groups: dict) -> bool:
"""Every row and every column hides none, or at least two.
Public so the tests can assert it directly rather than re-deriving it.
"""
return not _lone_hidden_groups(groups) and not _lone_hidden_categories(groups)
def _mask_for_disclosure(groups: dict) -> None:
"""Withhold further cells until nothing suppressed can be solved for.
Not rendering a figure is not the same as not publishing it. This endpoint
is public and unauthenticated, so anything left in the payload is
published, whatever the UI chooses to draw — the same reasoning the
admission_distance field carries in app.py.
Two identities let a caller solve for a withheld cell:
* within a pupil group, the categories sum to the cohort, so a group with
exactly ONE suppressed category gives it away as cohort - sum(rest);
* across groups, disadvantaged + other = all for every category, so a
category suppressed in exactly ONE of the three gives itself away.
DfE's own answer is secondary suppression: withhold a second cell so the
residual spans two unknowns and identifies neither.
Where no companion can do that — a sparse cohort whose every other category
is `not_applicable`, which is common in special schools and alternative
provision — there is nothing left to withhold, so the pupil group is
DROPPED entirely. An earlier version simply gave up here and returned with
the violation intact and no signal, which is the one outcome this function
must never produce: a disclosure-control pass that fails silently is worse
than none, because everything downstream trusts it.
Mutates `groups` in place. Guaranteed to return with
disclosure_invariant_holds(groups) true.
"""
def suppress(cell):
if cell["status"] == "published":
cell["status"] = "suppressed"
cell["pupils"] = None
cell["percentage"] = None
return True
return False
def add_companion(candidates) -> bool:
"""Withhold a second cell so the residual spans two unknowns.
The companion must carry pupils. Suppressing a zero looks like
secondary suppression and protects nothing: the residual still equals
the original withheld figure exactly. Returns False when no cell can
do the job, which escalates to dropping the group.
"""
published = [c for c in candidates if c["status"] == "published"]
useful = sorted(
(c for c in published if (c["pupils"] or 0) > 0),
key=lambda c: c["pupils"],
)
if useful:
return suppress(useful[0])
# Every remaining cell is zero or not applicable: withholding any of
# them leaves the residual equal to the original figure.
return False
# Fixpoint: each new suppression can break the other identity. Terminates
# because every pass either adds a suppression, drops a group, or stops.
while not disclosure_invariant_holds(groups):
changed = False
for category in _lone_hidden_categories(groups):
siblings = [
c for g in groups.values() for c in g["categories"]
if c["category"] == category
]
if add_companion(siblings):
changed = True
for key in _lone_hidden_groups(groups):
if add_companion(groups[key]["categories"]):
changed = True
if changed:
continue
# Nothing left to withhold. Drop the groups that are still solvable,
# and any category still solvable across the groups that remain.
for key in _lone_hidden_groups(groups):
del groups[key]
changed = True
for category in _lone_hidden_categories(groups):
for group in groups.values():
for cell in group["categories"]:
if cell["category"] == category and suppress(cell):
changed = True
if not changed:
# Unreachable given the two escalations above, but a masking pass
# must never spin or exit unsafely. Withhold everything.
groups.clear()
return
def _destinations_block(rows: list) -> dict | None:
"""Shape destination rows for one phase into the API's block.
Applies secondary suppression before returning, so no caller of this public
endpoint can solve for a figure DfE withheld. See _mask_for_disclosure.
Aggregate measures are dropped entirely. DfE publishes them, and they would
be useful for a "what is published for this group" fallback, but nothing
renders them today and an aggregate spanning exactly one suppressed
component names that component. An unused field that leaks is not a
trade-off worth carrying — re-add them with their own guard if the fallback
is ever built.
Deliberately computes no residual, no "remaining pupils" figure, and no
total that would close a gap left by a suppressed category.
"""
if not rows:
return None
years = [r["year"] for r in rows if r.get("year") is not None]
if not years:
return None
latest_year = max(years)
rows = [r for r in rows if r.get("year") == latest_year]
groups: dict = {}
for row in rows:
group = groups.setdefault(
row["pupil_group"],
{"cohort": row.get("cohort_pupils"), "categories": []},
)
measure = row["destination_measure"]
published = row.get("status") == "published"
# Belt and braces: percentage is derived from the same source cell as
# pupils, but publishing one without the other would hand back the
# cohort (pupils / percentage) and with it the residual.
cell = {
"category": measure,
"pupils": row.get("pupils") if published else None,
"percentage": row.get("percentage") if published else None,
"status": row.get("status"),
}
if measure in _AGGREGATE_MEASURES:
continue
group["categories"].append(cell)
if not groups:
return None
_mask_for_disclosure(groups)
# Masking can empty the block entirely — a sparse cohort where no group
# could be made safe. Return None so the section is absent rather than
# rendering an empty shell.
if not groups:
return None
return {"cohort_year": _format_cohort_year(latest_year), "groups": groups}
def _empty_supplementary() -> dict: def _empty_supplementary() -> dict:
return { return {
"ofsted": None, "ofsted": None,
@@ -775,6 +1040,7 @@ def _empty_supplementary() -> dict:
"phonics": None, "phonics": None,
"deprivation": None, "deprivation": None,
"finance": None, "finance": None,
"destinations": None,
} }
@@ -902,6 +1168,38 @@ def get_supplementary_data_batch(db: Session, urns: list[int]) -> dict:
result[f.urn]["finance"] = _finance_dict(f) result[f.urn]["finance"] = _finance_dict(f)
_safe(_finance) _safe(_finance)
# Destinations — KS4 and 16-18. Both marts are long-format, so every row
# for a URN is collected and _destinations_block picks the latest year and
# shapes the pupil groups. A phase with no rows serialises as null rather
# than an empty shell, so the frontend renders nothing rather than an empty
# section.
def _destinations():
from collections import defaultdict
def _collect(model):
per_urn = defaultdict(list)
for r in db.query(model).filter(model.urn.in_(urns)).all():
per_urn[r.urn].append({
"year": r.year,
"pupil_group": r.pupil_group,
"destination_measure": r.destination_measure,
"cohort_pupils": r.cohort_pupils,
"pupils": r.pupils,
"percentage": r.percentage,
"status": r.status,
})
return per_urn
ks4_rows = _collect(FactKs4Destinations)
ks5_rows = _collect(FactKs5Destinations)
for urn in urns:
ks4 = _destinations_block(ks4_rows.get(urn, []))
ks5 = _destinations_block(ks5_rows.get(urn, []))
result[urn]["destinations"] = (
{"ks4": ks4, "ks5": ks5} if (ks4 or ks5) else None
)
_safe(_destinations)
return result return result
+135
View File
@@ -0,0 +1,135 @@
"""Feature flags: what can be switched, and what is switched right now.
Ship-dark, not a kill switch. Flags let work merge and deploy without becoming
visible; they are expected to flip about monthly, by a person, deliberately.
Nothing here does percentage rollouts or user targeting — the site has no user
identity to target.
Unleash holds the state. It does not hold the list. REGISTRY below is that
list, and it exists for three reasons: the SDK evaluates an unknown flag to
False, so without a registry that is an *undeclared* False, indistinguishable
from a typo; /api/flags needs a key set to return when Unleash is unreachable;
and a flag in the UI but not in the registry is orphaned and should be visibly
so rather than quietly authoritative.
Every flag defaults to False. There is no per-flag default, because a flag that
defaults on is a kill switch, and this is not one.
"""
from __future__ import annotations
import logging
from dataclasses import dataclass
from datetime import date
from .config import settings
logger = logging.getLogger(__name__)
# A flag is temporary scaffolding. See test_a_flag_older_than_the_limit.
MAX_FLAG_AGE_DAYS = 90
@dataclass(frozen=True)
class Flag:
# One string: the registry key, the Unleash flag name, and the JSON key in
# /api/flags. snake_case, matching the API's existing convention. No case
# transformation anywhere, so there is no mapping layer to get wrong.
name: str
description: str # one line: what turning this on reveals
added: date # for the staleness tripwire
REGISTRY: dict[str, Flag] = {
f.name: f for f in (
Flag(
name="admission_distance",
description=(
"The last-distance-offered figure on the Admissions tile and "
"the 'How far away are you?' section on school pages."
),
added=date(2026, 8, 23),
),
Flag(
name="school_autosuggest",
description=(
"School name suggestions as you type in the main search box."
),
added=date(2026, 8, 26),
),
Flag(
name="about_page",
description=(
"The /about page, its footer link, its sitemap entry, and the "
"named-author byline on every blog post."
),
added=date(2026, 9, 8),
),
Flag(
name="blog",
description=(
"The /blog index, post pages, the RSS feed, their footer link "
"and their sitemap entries. Not /admin: posts must be "
"writable before the blog is readable."
),
added=date(2026, 9, 8),
),
)
}
_client = None
def init() -> None:
"""Start the Unleash client, or log why flags are all off.
Called once from the app lifespan. Never raises: a flag system that can
stop the API from booting is worse than one that is switched off.
"""
global _client
if not settings.unleash_url or not settings.unleash_api_token:
logger.warning(
"Unleash is not configured (UNLEASH_URL / UNLEASH_API_TOKEN); "
"every feature flag evaluates to False.")
return
try:
from UnleashClient import UnleashClient
_client = UnleashClient(
url=settings.unleash_url,
app_name=settings.unleash_app_name,
custom_headers={"Authorization": settings.unleash_api_token},
cache_directory=settings.unleash_cache_directory,
refresh_interval=15,
)
_client.initialize_client()
logger.info("Unleash client initialised against %s", settings.unleash_url)
except Exception:
# Fail closed and keep serving. The SDK also evaluates everything False
# until its first successful sync, so this is the same direction.
_client = None
logger.exception("Unleash client failed to start; flags are all False.")
def is_enabled(name: str) -> bool:
"""Whether `name` is on. False for anything unknown, unreachable or broken."""
if name not in REGISTRY:
logger.error(
"undeclared feature flag %r was evaluated; returning False. "
"Add it to backend/flags.py REGISTRY or fix the name.", name)
return False
if _client is None:
return False
try:
return bool(_client.is_enabled(
name, fallback_function=lambda feature_name, context: False))
except Exception:
logger.exception("flag %r failed to evaluate; returning False", name)
return False
def all_flags() -> dict[str, bool]:
"""Every declared flag and its current value. Serves /api/flags."""
return {name: is_enabled(name) for name in REGISTRY}
+45
View File
@@ -321,3 +321,48 @@ class Ks2NationalAverage(Base):
gps_high_pct = Column(Float) gps_high_pct = Column(Float)
gps_avg_score = Column(Float) gps_avg_score = Column(Float)
science_expected_pct = Column(Float) science_expected_pct = Column(Float)
class FactKs4Destinations(Base):
"""KS4 leavers destinations — one row per URN, year, pupil group, measure.
Long format rather than wide because pupil_group is a real third dimension.
`status` is load-bearing: 'suppressed' means DfE withheld a figure it
considered disclosive and the page must print "withheld"; 'not_applicable'
means the measure does not apply and the page must print nothing. `pupils`
is null for both, so collapsing status to a null check loses the
difference — and the categories sum to the cohort, so a consumer that
treats a withheld cell as zero republishes what DfE hid.
"""
__tablename__ = "fact_ks4_destinations"
__table_args__ = (
Index("ix_ks4_dest_urn_year", "urn", "year"),
MARTS,
)
urn = Column(Integer, primary_key=True)
year = Column(Integer, primary_key=True)
pupil_group = Column(String(20), primary_key=True)
destination_measure = Column(String(40), primary_key=True)
cohort_pupils = Column(Integer)
pupils = Column(Integer)
percentage = Column(Float)
status = Column(String(20))
class FactKs5Destinations(Base):
"""16-18 study leavers destinations — same grain as FactKs4Destinations."""
__tablename__ = "fact_ks5_destinations"
__table_args__ = (
Index("ix_ks5_dest_urn_year", "urn", "year"),
MARTS,
)
urn = Column(Integer, primary_key=True)
year = Column(Integer, primary_key=True)
pupil_group = Column(String(20), primary_key=True)
destination_measure = Column(String(40), primary_key=True)
cohort_pupils = Column(Integer)
pupils = Column(Integer)
percentage = Column(Float)
status = Column(String(20))
+132 -22
View File
@@ -30,6 +30,12 @@ class Place:
name: str name: str
urns: tuple[int, ...] urns: tuple[int, ...]
parent_authority: str | None # authority NAME, for the 301 target parent_authority: str | None # authority NAME, for the 301 target
# Every authority the place meaningfully sits in, largest first. A quarter
# of outcodes and a third of towns straddle a boundary — SW19 is mostly
# Merton but partly Wandsworth — so naming only one asserts something
# false. parent_authority stays single because a redirect needs one
# target; this is what the page shows.
authorities: tuple[tuple[str, int], ...] = ()
# URNs per phase, so the per-phase threshold can be applied without # URNs per phase, so the per-phase threshold can be applied without
# re-querying. A place with 30 primaries and 2 secondaries publishes a # re-querying. A place with 30 primaries and 2 secondaries publishes a
# primary variant and no secondary one. # primary variant and no secondary one.
@@ -53,57 +59,151 @@ def _publishable_urns(df) -> set[int]:
return set(df.loc[df[cols].notna().any(axis=1), "urn"].astype(int)) return set(df.loc[df[cols].notna().any(axis=1), "urn"].astype(int))
# The measure a phase page is built around. A page with no results in this
# column has nothing a list of school names does not already give.
_PHASE_METRIC = {
"primary": "rwm_expected_pct",
"secondary": "attainment_8_score",
}
def _phase_urns(group, publishable: set[int]) -> dict[str, tuple[int, ...]]: def _phase_urns(group, publishable: set[int]) -> dict[str, tuple[int, ...]]:
"""URNs per phase. All-through schools count toward both, matching the """URNs per phase, counting only schools with a result for that phase.
PHASE_GROUPS mapping the search filters already use."""
Not merely "publishable". A school with an Ofsted grade and no results is
worth a page of its own and belongs in the place list, but it cannot
populate a phase page's results column — and the threshold is there to ask
whether that column will have anything in it.
Counting publishable schools instead let /schools/kent/primary publish
with none of its five rows carrying a result, and left 44 phase pages
majority-blank. It is the same rule as "no page without a local average",
which was never extended per phase.
All-through schools count toward both phases, matching the PHASE_GROUPS
mapping the search filters already use.
"""
from backend.app import PHASE_GROUPS from backend.app import PHASE_GROUPS
if "phase" not in group.columns: if "phase" not in group.columns:
return {} return {}
lowered = group["phase"].fillna("").str.lower() lowered = group["phase"].fillna("").str.lower()
out: dict[str, tuple[int, ...]] = {} out: dict[str, tuple[int, ...]] = {}
for phase in ("primary", "secondary"): for phase in ("primary", "secondary"):
wanted = PHASE_GROUPS.get(phase, set()) wanted = PHASE_GROUPS.get(phase, set())
subset = group[lowered.isin(wanted)] subset = group[lowered.isin(wanted)]
# The page lists every school of the phase; the threshold counts only
# those carrying a result, so a mostly-empty table never publishes.
metric = _PHASE_METRIC[phase]
with_result = (
{int(u) for u in subset.loc[subset[metric].notna(), "urn"]}
if metric in subset.columns else set()
)
if len(with_result & publishable) < MIN_SCHOOLS:
continue
urns = tuple(sorted({int(u) for u in subset["urn"]} & publishable)) urns = tuple(sorted({int(u) for u in subset["urn"]} & publishable))
if urns: if urns:
out[phase] = urns out[phase] = urns
return out return out
def _parent_authority(group) -> str | None: # A place is described by an authority when it holds at least a tenth of the
"""The most common authority in a group — the useful 301 target. # schools, and at least two. GIAS carries occasional postcode errors — EN6
# lists two Shropshire schools among fourteen in Hertfordshire — and a bare
# "any authority present" rule would print those as though they were real.
# There is deliberately no cap on how many are named. An earlier cut stopped
# at three, which silently dropped the fourth in exactly the case where the
# information matters most — a genuinely fragmented place. The share rule is
# the only limit, and it already bounds the list at ten.
_AUTHORITY_MIN_SHARE = 0.10
_AUTHORITY_MIN_SCHOOLS = 2
def _authorities(group) -> tuple[tuple[str, int], ...]:
"""Authorities this place meaningfully sits in, largest first."""
from backend.app import EXCLUDED_FILTER_VALUES
A town spanning several authorities has no single parent, so the mode is
the honest answer rather than an arbitrary first row.
"""
if "local_authority" not in group.columns: if "local_authority" not in group.columns:
return None return ()
top = group["local_authority"].dropna() counts = group["local_authority"].dropna().value_counts()
return str(top.mode().iloc[0]) if not top.empty else None total = int(counts.sum())
if not total:
return ()
kept = [
(str(name), int(n)) for name, n in counts.items()
if str(name) not in EXCLUDED_FILTER_VALUES
and n >= _AUTHORITY_MIN_SCHOOLS
and n / total >= _AUTHORITY_MIN_SHARE
]
# A place too small or too fragmented for the share rule still names its
# largest authority, or the page would say nothing about where it is.
if not kept:
for name, n in counts.items():
if str(name) not in EXCLUDED_FILTER_VALUES:
return ((str(name), int(n)),)
return ()
return tuple(kept)
def _parent_authority(authorities: tuple[tuple[str, int], ...]) -> str | None:
"""The 301 target: the largest authority a place sits in.
Derived from `authorities` rather than computed separately. The first cut
used `mode()` here while `authorities` used `value_counts()`, and on an
exact tie pandas does not guarantee the two pick the same name — so the
redirect could have pointed somewhere other than the authority the page
named first. One computation, one answer.
Deriving it also inherits the sentinel filter, so a place can no longer
redirect to /schools/authority/does-not-apply.
"""
return authorities[0][0] if authorities else None
def _group(df, column: str, kind: str, publishable: set[int]) -> dict[str, Place]: def _group(df, column: str, kind: str, publishable: set[int]) -> dict[str, Place]:
"""One Place per distinct value of `column` that clears the threshold.""" """One Place per distinct SLUG in `column` that clears the threshold.
Grouped by slug, not by raw value, because GIAS spells the same place
several ways and they all resolve to one URL. Five town slugs come from
more than one spelling: "London" (1,819 schools) and "LONDON" (12) both
slugify to `london`; Weston-super-Mare is split 14/19 across two
spellings; Newcastle-under-Lyme across three.
Grouping by raw value meant the later group simply overwrote the earlier
one in this dict — so /schools/london could have shown twelve schools
instead of 1,819, silently and depending on row order.
The display name is the most common spelling, which is the one a reader
expects to see.
"""
from backend.app import _slugify from backend.app import _slugify
if column not in df.columns: if column not in df.columns:
return {} return {}
working = df.assign(_slug=df[column].map(
lambda v: _slugify(str(v).strip()) if isinstance(v, str) and v.strip() else None))
working = working[working["_slug"].notna() & (working["_slug"] != "")]
out: dict[str, Place] = {} out: dict[str, Place] = {}
for name, group in df.groupby(column, dropna=True): for slug, group in working.groupby("_slug"):
name = str(name).strip() slug = str(slug)
if not name:
continue
urns = tuple(sorted({int(u) for u in group["urn"]} & publishable)) urns = tuple(sorted({int(u) for u in group["urn"]} & publishable))
if len(urns) < MIN_SCHOOLS: if len(urns) < MIN_SCHOOLS:
continue continue
slug = _slugify(name) spellings = group[column].dropna().value_counts()
if not slug: if spellings.empty:
continue continue
name = str(spellings.index[0]).strip()
authorities = () if kind == "authority" else _authorities(group)
place = Place( place = Place(
kind=kind, slug=slug, name=name, urns=urns, kind=kind, slug=slug, name=name, urns=urns,
parent_authority=_parent_authority(group) if kind == "town" else None, parent_authority=_parent_authority(authorities),
authorities=authorities,
phase_urns=_phase_urns(group, publishable), phase_urns=_phase_urns(group, publishable),
) )
out[place.key] = place out[place.key] = place
@@ -124,7 +224,14 @@ def _outcode(postcode) -> str | None:
def _outcode_places(df, publishable: set[int]) -> dict[str, Place]: def _outcode_places(df, publishable: set[int]) -> dict[str, Place]:
"""One Place per postcode district clearing the threshold. """One Place per postcode district clearing the threshold.
These carry no phase variants: nobody searches "primary schools in SW11". These carry no phase variants: nobody searches "primary schools in SW11",
so the spec gives them no /primary or /secondary route. `phase_urns` is
left empty rather than computed and then filtered downstream — the
registry is the one place that decides which phases a place publishes,
and the page links whatever it reports.
Computing them here put a link to a route that does not exist on every one
of the 1,720 outcode pages.
""" """
if "postcode" not in df.columns: if "postcode" not in df.columns:
return {} return {}
@@ -136,9 +243,10 @@ def _outcode_places(df, publishable: set[int]) -> dict[str, Place]:
urns = tuple(sorted({int(u) for u in group["urn"]} & publishable)) urns = tuple(sorted({int(u) for u in group["urn"]} & publishable))
if len(urns) < MIN_SCHOOLS: if len(urns) < MIN_SCHOOLS:
continue continue
authorities = _authorities(group)
place = Place(kind="outcode", slug=str(oc).lower(), name=str(oc), place = Place(kind="outcode", slug=str(oc).lower(), name=str(oc),
urns=urns, parent_authority=_parent_authority(group), urns=urns, parent_authority=_parent_authority(authorities),
phase_urns=_phase_urns(group, publishable)) authorities=authorities)
out[place.key] = place out[place.key] = place
return out return out
@@ -179,8 +287,10 @@ def _locality_places(df, publishable: set[int],
"threshold of %d - not published", "threshold of %d - not published",
slug, ", ".join(outcodes), len(urns), MIN_SCHOOLS) slug, ", ".join(outcodes), len(urns), MIN_SCHOOLS)
continue continue
authorities = _authorities(group)
place = Place(kind="locality", slug=slug, name=name, urns=urns, place = Place(kind="locality", slug=slug, name=name, urns=urns,
parent_authority=_parent_authority(group), parent_authority=_parent_authority(authorities),
authorities=authorities,
phase_urns=_phase_urns(group, publishable)) phase_urns=_phase_urns(group, publishable))
out[place.key] = place out[place.key] = place
return out return out
+269
View File
@@ -0,0 +1,269 @@
"""The destinations serialiser's contract.
Not rendering a figure is not the same as not publishing it. This endpoint is
public and unauthenticated, so whatever the payload carries is published,
whatever the UI draws. The categories sum to the cohort and the pupil groups
sum to each other, so a lone suppressed cell is solvable by subtraction — the
serialiser adds secondary suppression to prevent it.
See docs/superpowers/specs/2026-08-28-destination-measures-design.md.
"""
from backend.data_loader import (
_destinations_block, _format_cohort_year, disclosure_invariant_holds,
)
def _row(group, measure, pupils, status, cohort=180, percentage=None, year=202223):
return {
"pupil_group": group,
"destination_measure": measure,
"pupils": pupils,
"percentage": percentage,
"status": status,
"cohort_pupils": cohort,
"year": year,
}
def test_suppressed_category_serialises_as_suppressed_with_null_pupils():
rows = [
_row("all", "school_sixth_form", 75, "published", percentage=41.7),
_row("all", "sixth_form_college", None, "suppressed"),
]
block = _destinations_block(rows)
cats = {c["category"]: c for c in block["groups"]["all"]["categories"]}
assert cats["sixth_form_college"]["status"] == "suppressed"
assert cats["sixth_form_college"]["pupils"] is None
assert cats["sixth_form_college"]["percentage"] is None
def test_published_category_keeps_its_figures():
block = _destinations_block([
_row("all", "school_sixth_form", 75, "published", percentage=41.7),
])
cat = block["groups"]["all"]["categories"][0]
assert cat["pupils"] == 75
assert cat["percentage"] == 41.7
assert cat["status"] == "published"
def test_only_the_latest_year_is_served():
rows = [
_row("all", "school_sixth_form", 60, "published", year=202122),
_row("all", "school_sixth_form", 75, "published", year=202223),
]
block = _destinations_block(rows)
assert block["cohort_year"] == "2022/23"
assert len(block["groups"]["all"]["categories"]) == 1
assert block["groups"]["all"]["categories"][0]["pupils"] == 75
def test_all_three_pupil_groups_are_carried():
rows = [
_row("all", "school_sixth_form", 75, "published"),
_row("disadvantaged", "school_sixth_form", 17, "published", cohort=62),
_row("other", "school_sixth_form", 58, "published", cohort=118),
]
block = _destinations_block(rows)
assert set(block["groups"]) == {"all", "disadvantaged", "other"}
assert block["groups"]["disadvantaged"]["cohort"] == 62
def test_cohort_year_is_reported_so_the_page_can_date_itself():
block = _destinations_block([_row("all", "school_sixth_form", 75, "published")])
assert block["cohort_year"] == "2022/23"
def test_format_cohort_year_handles_the_six_digit_form():
assert _format_cohort_year(202223) == "2022/23"
assert _format_cohort_year(None) is None
def test_empty_rows_yield_none_not_an_empty_shell():
assert _destinations_block([]) is None
# ── Disclosure control ──────────────────────────────────────────────────────
#
# The rendering guards in lib/destinations.ts stop a withheld figure being
# DRAWN. They do nothing about it being COMPUTED: this endpoint is public and
# unauthenticated, so whatever the payload carries is published. These tests
# are the ones that matter.
def _solve_residual(group):
"""What any caller can work out: cohort minus everything published."""
published = [c["pupils"] for c in group["categories"] if c["pupils"] is not None]
hidden = [c for c in group["categories"] if c["status"] == "suppressed"]
return group["cohort"] - sum(published), len(hidden)
def test_a_lone_suppressed_category_cannot_be_solved_for():
"""Whitley Bay High School's real 2022/23 disadvantaged group: further
education withheld, everything else published, cohort 41. Before secondary
suppression the payload gave the answer away as 41 - 23 = 18."""
rows = [
_row("disadvantaged", "school_sixth_form", 15, "published", cohort=41),
_row("disadvantaged", "sixth_form_college", 0, "published", cohort=41),
_row("disadvantaged", "further_education", None, "suppressed", cohort=41),
_row("disadvantaged", "apprenticeship", 1, "published", cohort=41),
_row("disadvantaged", "employment", 2, "published", cohort=41),
_row("disadvantaged", "not_sustained", 3, "published", cohort=41),
_row("disadvantaged", "not_captured", 2, "published", cohort=41),
]
group = _destinations_block(rows)["groups"]["disadvantaged"]
residual, hidden = _solve_residual(group)
assert hidden >= 2, "a lone suppressed cell must gain a companion"
assert residual != 18, "the withheld figure is recoverable from the payload"
def test_every_group_hides_none_or_at_least_two_categories():
rows = [
_row("all", "school_sixth_form", 75, "published"),
_row("all", "sixth_form_college", None, "suppressed"),
_row("all", "further_education", 61, "published"),
_row("all", "apprenticeship", 8, "published"),
_row("all", "employment", 6, "published"),
_row("all", "not_sustained", 5, "published"),
_row("all", "not_captured", 4, "published"),
]
group = _destinations_block(rows)["groups"]["all"]
hidden = [c for c in group["categories"] if c["status"] == "suppressed"]
assert len(hidden) >= 2
def test_a_category_hidden_in_one_group_is_hidden_in_a_second():
"""disadvantaged + other = all for every category, so a category withheld
in exactly one of the three is recoverable from the other two."""
rows = []
for measure, a, d, o in [
("school_sixth_form", 75, None, 58),
("further_education", 61, 27, 34),
("apprenticeship", 8, 4, 4),
("employment", 6, 1, 5),
("not_sustained", 5, 3, 2),
("not_captured", 4, 2, 2),
]:
rows.append(_row("all", measure, a, "published", cohort=159))
rows.append(_row("disadvantaged", measure, d,
"published" if d is not None else "suppressed", cohort=37))
rows.append(_row("other", measure, o, "published", cohort=122))
groups = _destinations_block(rows)["groups"]
measures = {c["category"] for g in groups.values() for c in g["categories"]}
assert len(measures) == 6, "the fixture's six measures must all be checked"
for measure in sorted(measures):
hidden = sum(
1 for g in groups.values() for c in g["categories"]
if c["category"] == measure and c["status"] == "suppressed"
)
# The invariant is "none, or at least two" — not "at least two".
assert hidden != 1, f"{measure} is solvable across the pupil groups"
def test_a_suppressed_cell_never_keeps_its_percentage():
"""percentage / pupils would hand back the cohort, and with it the residual."""
rows = [
_row("all", "school_sixth_form", 75, "published", percentage=41.7),
_row("all", "sixth_form_college", None, "suppressed", percentage=11.7),
_row("all", "further_education", 61, "published", percentage=33.9),
]
group = _destinations_block(rows)["groups"]["all"]
for cell in group["categories"]:
if cell["status"] != "published":
assert cell["pupils"] is None
assert cell["percentage"] is None
def test_aggregates_are_not_served():
"""An aggregate spanning exactly one suppressed component names it, and
nothing renders them today."""
rows = [
_row("all", "school_sixth_form", 75, "published"),
_row("all", "agg_sustained_all", 171, "published"),
]
group = _destinations_block(rows)["groups"]["all"]
assert [c["category"] for c in group["categories"]] == ["school_sixth_form"]
assert "aggregates" not in group
def test_a_fully_published_group_is_left_alone():
"""Secondary suppression must not cost anything where nothing is withheld —
this is the all-pupils view on every mainstream secondary."""
rows = [
_row("all", m, p, "published")
for m, p in [("school_sixth_form", 75), ("sixth_form_college", 21),
("further_education", 61), ("apprenticeship", 8),
("employment", 6), ("not_sustained", 5), ("not_captured", 4)]
]
group = _destinations_block(rows)["groups"]["all"]
assert all(c["status"] == "published" for c in group["categories"])
assert len(group["categories"]) == 7
def test_the_invariant_is_asserted_directly_not_re_derived():
"""A group with one suppressed category and nothing else to withhold."""
rows = [
_row("all", "school_sixth_form", None, "suppressed", cohort=9),
_row("all", "sixth_form_college", None, "not_applicable", cohort=9),
_row("all", "further_education", None, "not_applicable", cohort=9),
]
block = _destinations_block(rows)
assert block is None or disclosure_invariant_holds(block["groups"])
def test_a_sparse_cohort_with_no_companion_drops_the_group():
"""Special schools and AP routinely have one suppressed category and every
other one not applicable. There is nothing left to withhold, so the group
goes — an earlier version returned here with the violation intact."""
rows = [
_row("all", "school_sixth_form", None, "suppressed", cohort=9),
_row("all", "sixth_form_college", None, "not_applicable", cohort=9),
_row("all", "further_education", None, "not_applicable", cohort=9),
_row("all", "apprenticeship", None, "not_applicable", cohort=9),
_row("all", "employment", None, "not_applicable", cohort=9),
_row("all", "not_sustained", None, "not_applicable", cohort=9),
_row("all", "not_captured", None, "not_applicable", cohort=9),
]
block = _destinations_block(rows)
assert block is None or "all" not in block["groups"], (
"a group that cannot be made safe must not be served"
)
def test_zeros_are_not_treated_as_a_usable_companion():
"""Suppressing a zero protects nothing — the residual is unchanged. With
only zeros available the group must be dropped, not falsely 'fixed'."""
rows = [
_row("all", "school_sixth_form", None, "suppressed", cohort=5),
_row("all", "sixth_form_college", 0, "published", cohort=5),
_row("all", "further_education", 0, "published", cohort=5),
]
block = _destinations_block(rows)
if block and "all" in block["groups"]:
group = block["groups"]["all"]
published = sum(c["pupils"] for c in group["categories"]
if c["pupils"] is not None)
hidden = [c for c in group["categories"] if c["status"] == "suppressed"]
assert len(hidden) != 1, "a zero companion leaves the figure solvable"
assert group["cohort"] - published != 5
def test_masking_always_terminates_in_a_safe_state():
"""Exhaustive over every suppression pattern of a four-category group."""
from itertools import product
MEASURES = ["school_sixth_form", "sixth_form_college",
"further_education", "apprenticeship"]
for statuses in product(["published", "suppressed", "not_applicable"],
repeat=len(MEASURES)):
rows = [
_row("all", m, 3 if st == "published" else None, st, cohort=12)
for m, st in zip(MEASURES, statuses)
]
block = _destinations_block(rows)
if block is None:
continue
assert disclosure_invariant_holds(block["groups"]), (
f"invariant broken for {statuses}"
)
+163
View File
@@ -0,0 +1,163 @@
"""Tests for the feature flag layer (spec 2026-08-23).
None of these need a running Unleash. That is the point: an unset UNLEASH_URL
means every flag is False, which is what local development and CI get.
"""
from datetime import date, timedelta
from backend import flags
def test_every_declared_flag_is_keyed_by_its_own_name():
# One string is the registry key, the Unleash flag name and the JSON key.
# A mismatch here would mean the UI toggles a flag the code never reads.
for key, flag in flags.REGISTRY.items():
assert key == flag.name
def test_flag_names_are_snake_case():
# Matches the API's existing convention (admission_distance,
# rwm_expected_pct) so no case transformation exists to get wrong.
for name in flags.REGISTRY:
assert name == name.lower()
assert "-" not in name and " " not in name
def test_an_unconfigured_client_evaluates_every_flag_false(monkeypatch):
monkeypatch.setattr(flags, "_client", None)
for name in flags.REGISTRY:
assert flags.is_enabled(name) is False
def test_an_undeclared_flag_is_false_rather_than_an_error(monkeypatch):
# A typo'd flag name must not raise in a request path. It is logged as an
# error, because an undeclared flag is always a bug.
monkeypatch.setattr(flags, "_client", None)
assert flags.is_enabled("no_such_flag") is False
def test_an_exploding_client_is_false_rather_than_a_500(monkeypatch):
class Boom:
def is_enabled(self, *a, **kw):
raise RuntimeError("unleash is on fire")
monkeypatch.setattr(flags, "_client", Boom())
name = next(iter(flags.REGISTRY))
assert flags.is_enabled(name) is False
def test_all_flags_reports_every_declared_flag(monkeypatch):
monkeypatch.setattr(flags, "_client", None)
assert set(flags.all_flags()) == set(flags.REGISTRY)
assert all(v is False for v in flags.all_flags().values())
def test_a_flag_older_than_the_limit_fails_this_test():
"""A tripwire, not an assertion about correctness.
Flags are temporary scaffolding and the failure mode of every flag system
is accumulation. This fails on the day a flag turns 90, on whatever PR
happens to be open — which is the point: someone has to decide.
To fix: delete the flag and the branches that read it, or, if it genuinely
still needs to exist, move its `added` date and say why in the commit.
"""
stale = [
f.name for f in flags.REGISTRY.values()
if date.today() - f.added > timedelta(days=flags.MAX_FLAG_AGE_DAYS)
]
assert not stale, (
f"Flags older than {flags.MAX_FLAG_AGE_DAYS} days: {stale}. "
"Remove the flag and the code branches it guards, or move its `added` "
"date deliberately."
)
def _client():
from fastapi.testclient import TestClient
from backend import app as app_module
return TestClient(app_module.app, raise_server_exceptions=False)
def test_the_flags_endpoint_lists_every_declared_flag(monkeypatch):
monkeypatch.setattr(flags, "_client", None)
body = _client().get("/api/flags").json()
assert set(body) == set(flags.REGISTRY)
def test_the_flags_endpoint_answers_false_when_unleash_is_unreachable(monkeypatch):
# The endpoint must still answer. A frontend that cannot read flags renders
# everything dark, which is right; one that gets a 500 renders nothing.
monkeypatch.setattr(flags, "_client", None)
res = _client().get("/api/flags")
assert res.status_code == 200
assert all(v is False for v in res.json().values())
def _school_payload(monkeypatch, *, flag_on: bool):
"""Fetch one school's payload with the distance flag forced on or off.
The DataFrame shape is copied from test_school_details.py rather than
minimised: the endpoint reads a wide set of GIAS columns, and a trimmed
frame fails for reasons that have nothing to do with flags.
"""
import numpy as np
import pandas as pd
from fastapi.testclient import TestClient
from backend import app as app_module
df = pd.DataFrame([{
"urn": 150275,
"school_name": "West London Performing Arts Academy",
"phase": "Secondary",
"school_type": "Special post 16 institution",
"trust_name": None,
"religious_denomination": "Does not apply",
"gender": None,
"age_range": "16-25",
"admissions_policy": None,
"capacity": np.nan,
"gias_total_pupils": np.nan,
"headteacher_name": None,
"website": None,
"ofsted_grade": np.nan,
"local_authority": "Ealing",
"address": "268 Northfield Avenue, London, W5 4UB",
"postcode": "W5 4UB",
"latitude": 51.4986,
"longitude": -0.3148,
"year": np.nan,
"total_pupils": np.nan,
"eligible_pupils": np.nan,
"rwm_expected_pct": np.nan,
}])
monkeypatch.setattr(app_module, "load_school_data", lambda: df)
# Two arguments: get_supplementary_data(db, urn). See backend/app.py.
monkeypatch.setattr(
app_module, "get_supplementary_data",
lambda db, urn: {"admission_distance": {"distance_m": 772.49,
"year": 2024}})
monkeypatch.setattr(flags, "is_enabled", lambda name: flag_on)
client = TestClient(app_module.app, raise_server_exceptions=False)
res = client.get("/api/schools/150275")
assert res.status_code == 200, res.text
return res.json()
def test_the_distance_field_is_absent_when_the_flag_is_off(monkeypatch):
"""Absent, not null, and withheld at the source.
/api/schools/ is public and unauthenticated. Leaving a withheld field in
the payload while declining to render it hands the record to anyone who
opens the network tab — the reasoning already recorded in c9a1892.
"""
body = _school_payload(monkeypatch, flag_on=False)
assert "admission_distance" not in body
def test_the_distance_field_is_present_when_the_flag_is_on(monkeypatch):
body = _school_payload(monkeypatch, flag_on=True)
assert body["admission_distance"]["distance_m"] == 772.49
+189
View File
@@ -229,3 +229,192 @@ def test_no_curated_locality_names_a_london_borough():
f"these are boroughs, not districts: {sorted(named)} - they already " f"these are boroughs, not districts: {sorted(named)} - they already "
"have an authority page covering every school" "have an authority page covering every school"
) )
def test_a_place_names_every_authority_it_straddles():
"""SW19 is mostly Merton but partly Wandsworth.
A quarter of viable outcodes and a third of viable towns cross an
authority boundary, so naming only the largest asserts something false.
"""
rows = (_town(26, "London", "Merton", start=300000)
+ _town(7, "London", "Wandsworth", start=400000))
for r in rows:
r["postcode"] = "SW19 1AA"
reg = build_place_registry(_df(rows))
names = [n for n, _ in reg["outcode:sw19"].authorities]
assert names == ["Merton", "Wandsworth"] # largest first
assert dict(reg["outcode:sw19"].authorities)["Wandsworth"] == 7
def test_the_redirect_target_stays_a_single_authority():
# parent_authority and authorities do different jobs: a 301 needs one
# target, the page needs the truth.
rows = (_town(26, "London", "Merton", start=300000)
+ _town(7, "London", "Wandsworth", start=400000))
for r in rows:
r["postcode"] = "SW19 1AA"
reg = build_place_registry(_df(rows))
assert reg["outcode:sw19"].parent_authority == "Merton"
def test_a_stray_authority_below_the_share_threshold_is_not_named():
# GIAS carries postcode errors — EN6 lists two Shropshire schools among
# fourteen in Hertfordshire. Printing those as though real would be worse
# than omitting them.
rows = (_town(30, "Barnet", "Hertfordshire", start=300000)
+ _town(1, "Barnet", "Shropshire", start=400000))
for r in rows:
r["postcode"] = "EN6 1AA"
reg = build_place_registry(_df(rows))
assert [n for n, _ in reg["outcode:en6"].authorities] == ["Hertfordshire"]
def test_a_sentinel_authority_is_never_named():
rows = (_town(20, "London", "Merton", start=300000)
+ _town(6, "London", "Does not apply", start=400000))
for r in rows:
r["postcode"] = "SW19 1AA"
reg = build_place_registry(_df(rows))
assert [n for n, _ in reg["outcode:sw19"].authorities] == ["Merton"]
def test_a_place_always_names_at_least_one_authority():
# Even when every authority is below the share threshold, the page has to
# say where the place is.
rows = []
for i, la in enumerate(["A", "B", "C", "D", "E", "F", "G"]):
rows += _town(1, "Fragmented", la, start=300000 + i * 100)
reg = build_place_registry(_df(rows))
place = reg.get("town:fragmented")
assert place is not None
assert len(place.authorities) == 1
def test_every_qualifying_authority_is_named_with_no_cap():
"""An earlier cut stopped at three, dropping the fourth silently.
That truncation bit exactly where the information matters most — a
genuinely fragmented place — and nothing recorded it.
"""
rows = []
for i, la in enumerate(["Hackney", "Lambeth", "Westminster", "Lewisham"]):
rows += _town(3, "Fourway", la, start=300000 + i * 100)
reg = build_place_registry(_df(rows))
assert len(reg["town:fourway"].authorities) == 4
def test_the_redirect_target_is_the_authority_named_first():
"""They were computed separately — mode() against value_counts() — and on
an exact tie pandas does not guarantee the two agree."""
rows = (_town(26, "London", "Merton", start=300000)
+ _town(7, "London", "Wandsworth", start=400000))
for r in rows:
r["postcode"] = "SW19 1AA"
place = build_place_registry(_df(rows))["outcode:sw19"]
assert place.parent_authority == place.authorities[0][0]
def test_a_place_never_redirects_to_a_sentinel_authority():
# Deriving the parent from `authorities` inherits its sentinel filter.
rows = (_town(6, "Someplace", "Does not apply", start=300000)
+ _town(5, "Someplace", "Essex", start=400000))
reg = build_place_registry(_df(rows))
assert reg["town:someplace"].parent_authority == "Essex"
def test_spellings_of_one_place_are_merged_not_overwritten():
"""GIAS spells the same place several ways, and they share a URL.
"London" (1,819 schools) and "LONDON" (12) both slugify to `london`.
Grouping by raw value let the later group overwrite the earlier one, so
the page could have shown twelve schools instead of 1,819 — silently, and
depending on row order.
"""
rows = (_town(6, "Weston-super-Mare", "North Somerset", start=300000)
+ _town(5, "Weston-Super-Mare", "North Somerset", start=400000))
reg = build_place_registry(_df(rows))
assert len(reg["town:weston-super-mare"].urns) == 11
def test_the_merged_place_takes_its_most_common_spelling():
rows = (_town(9, "Newcastle-under-Lyme", "Staffordshire", start=300000)
+ _town(5, "NEWCASTLE-UNDER-LYME", "Staffordshire", start=400000))
reg = build_place_registry(_df(rows))
assert reg["town:newcastle-under-lyme"].name == "Newcastle-under-Lyme"
def test_a_phase_page_needs_results_not_merely_publishable_schools():
"""/schools/kent/primary published with none of its five rows scored.
The threshold counted schools that were publishable — a result OR an
Ofsted grade — while the page exists for its results column. Forty-four
phase pages were majority-blank; one had no results at all.
"""
rows = _town(MIN_SCHOOLS, "Kent", "Kent")
for r in rows:
r["rwm_expected_pct"] = np.nan # Ofsted only, no results
reg = build_place_registry(_df(rows))
assert "town:kent" in reg # the place still publishes
assert not reg["town:kent"].publishes_phase("primary")
def test_a_phase_page_publishes_once_enough_schools_carry_a_result():
rows = _town(MIN_SCHOOLS, "Beccles", "Suffolk")
reg = build_place_registry(_df(rows))
assert reg["town:beccles"].publishes_phase("primary")
def test_a_publishing_phase_page_still_lists_its_unscored_schools():
"""The threshold gates whether the page exists; it does not filter rows.
A parent looking up a school by name has to find it whether or not it
published results.
"""
scored = _town(MIN_SCHOOLS, "Beccles", "Suffolk", start=300000)
unscored = _town(2, "Beccles", "Suffolk", start=400000)
for r in unscored:
r["rwm_expected_pct"] = np.nan
reg = build_place_registry(_df(scored + unscored))
place = reg["town:beccles"]
assert place.publishes_phase("primary")
assert len(place.phase_urns["primary"]) == MIN_SCHOOLS + 2
def test_the_secondary_threshold_counts_its_own_metric():
# A town full of scored primaries must not thereby publish a secondary page.
rows = _town(MIN_SCHOOLS, "Brentwood", "Essex")
reg = build_place_registry(_df(rows))
assert not reg["town:brentwood"].publishes_phase("secondary")
def test_an_outcode_publishes_no_phase_variants():
"""There is no /schools/near/[outcode]/[phase] route, by design.
Nobody searches "primary schools in SW11", so the spec gives outcodes no
phase variants. The registry computed them anyway, and the place page —
which links whatever phases the registry reports — put two 404s on every
outcode page in the site.
This is the single rule now: a kind with no phase route reports no phases,
so neither the page nor the sitemap can offer one.
"""
rows = [{"urn": 500000 + i, "school_name": f"SW11 School {i}",
"town": "London", "local_authority": "Wandsworth",
"postcode": "SW11 1AA"} for i in range(MIN_SCHOOLS + 3)]
reg = build_place_registry(_df(rows))
place = reg["outcode:sw11"]
assert place.phase_urns == {}
assert not place.publishes_phase("primary")
assert not place.publishes_phase("secondary")
def test_an_authority_still_publishes_phase_variants():
"""Authorities keep theirs — "primary schools in Kent" is a real query,
and /schools/authority/[la]/[phase] is the route that serves it."""
reg = build_place_registry(_df(_town(MIN_SCHOOLS, "Maidstone", "Kent")))
assert reg["authority:kent"].publishes_phase("primary")
+116 -2
View File
@@ -45,10 +45,30 @@ def test_registry_carries_a_count_per_place(client):
assert town["count"] == 6 assert town["count"] == 6
def test_place_detail_returns_its_schools_ranked(client): def test_place_detail_returns_its_schools_alphabetically(client):
"""A place page is read by someone looking for a school they can name.
Scanning for it is what the order should serve, so the list is A-Z.
/api/rankings is where the league-table ordering lives.
"""
body = client.get("/api/places/town/brentwood").json() body = client.get("/api/places/town/brentwood").json()
assert body["place"]["name"] == "Brentwood" assert body["place"]["name"] == "Brentwood"
scores = [s["rwm_expected_pct"] for s in body["schools"]] names = [s["school_name"] for s in body["schools"]]
assert names == sorted(names, key=str.lower)
def test_place_ordering_ignores_case(client):
body = client.get("/api/places/town/brentwood").json()
names = [s["school_name"] for s in body["schools"]]
# A capitalised name must not sort ahead of every lowercase one.
assert names == sorted(names, key=str.lower)
def test_the_rankings_endpoint_still_ranks_by_metric(client):
# Alphabetical is a place-page decision, not a site-wide one.
body = client.get("/api/rankings?metric=rwm_expected_pct&phase=primary").json()
scores = [r["rwm_expected_pct"] for r in body.get("rankings", [])
if r.get("rwm_expected_pct") is not None]
assert scores == sorted(scores, reverse=True) assert scores == sorted(scores, reverse=True)
@@ -69,3 +89,97 @@ def test_unknown_place_404s(client):
def test_unknown_kind_404s(client): def test_unknown_kind_404s(client):
assert client.get("/api/places/planet/mars").status_code == 404 assert client.get("/api/places/planet/mars").status_code == 404
def _straddling_df() -> pd.DataFrame:
"""Eight schools in CM13: six in Essex, which has a page, and two in an
authority too small to have one.
Two, not one: the registry ignores an authority holding a single school in
a place, because GIAS carries occasional postcode errors."""
df = _schools_df()
extra = df.iloc[:2].copy()
extra["urn"] = [200000, 200001]
extra["school_name"] = ["Scilly School 0", "Scilly School 1"]
extra["local_authority"] = "Isles Of Scilly"
return pd.concat([df, extra], ignore_index=True)
@pytest.fixture()
def straddling_client(monkeypatch):
from backend import app as app_module
monkeypatch.setattr(app_module, "load_school_data", _straddling_df)
monkeypatch.setattr(app_module, "load_latest_school_data", _straddling_df)
monkeypatch.setattr(app_module, "_place_registry", None)
return TestClient(app_module.app, raise_server_exceptions=False)
def test_an_outcode_reports_no_phases_because_it_has_no_phase_route(client):
body = client.get("/api/places/outcode/cm13").json()
assert body["place"]["phases"] == []
def test_an_authority_reports_the_phases_it_publishes(client):
body = client.get("/api/places/authority/essex").json()
assert body["place"]["phases"] == ["primary"]
def test_an_authority_without_a_page_is_named_but_carries_no_slug(straddling_client):
"""Two English authorities — City of London and the Isles of Scilly — hold
fewer than the five schools a page needs, so they have no page.
Naming them is still right: the page says where the place is. Linking them
would not be. A null slug is what tells the page to print the name plainly
rather than invent a URL that 404s.
"""
body = straddling_client.get("/api/places/outcode/cm13").json()
by_name = {a["name"]: a for a in body["place"]["authorities"]}
assert by_name["Essex"]["slug"] == "essex"
assert by_name["Isles Of Scilly"]["slug"] is None
def _attributed_df() -> pd.DataFrame:
"""The same town, with the four attributes the place table now shows."""
df = _schools_df()
df["age_range"] = "4-11"
df["religious_denomination"] = "Church of England"
df["nursery_provision"] = True
df["parliamentary_constituency"] = "Brentwood and Ongar"
return df
@pytest.fixture()
def attributed_client(monkeypatch):
from backend import app as app_module
monkeypatch.setattr(app_module, "load_school_data", _attributed_df)
monkeypatch.setattr(app_module, "load_latest_school_data", _attributed_df)
monkeypatch.setattr(app_module, "_place_registry", None)
return TestClient(app_module.app, raise_server_exceptions=False)
def test_place_detail_carries_the_attributes_the_table_shows(attributed_client):
"""age_range and religious_denomination ride in on SCHOOL_COLUMNS.
nursery_provision and parliamentary_constituency do not, and the place
table needs all four — a column the response cannot fill is a column of
dashes on ~3,900 pages.
"""
body = attributed_client.get("/api/places/town/brentwood").json()
school = body["schools"][0]
assert school["age_range"] == "4-11"
assert school["religious_denomination"] == "Church of England"
assert school["nursery_provision"] is True
assert school["parliamentary_constituency"] == "Brentwood and Ongar"
def test_place_detail_survives_a_mart_without_the_optional_columns(client):
"""The base fixture has neither column, as an unrebuilt mart does not.
data_loader degrades those to NULL rather than failing the load, so the
endpoint must not assume they are present.
"""
res = client.get("/api/places/town/brentwood")
assert res.status_code == 200
assert "nursery_provision" not in res.json()["schools"][0]
+147
View File
@@ -0,0 +1,147 @@
"""The rate-limit bucket must be the caller, not the proxy in front of them.
`get_remote_address` reads request.client.host. In staging and production the
backend has no published ports and its only caller is the Next proxy, so that
host is the Next container — one bucket for every browser user on the site.
Measured before this fix: 70 concurrent requests, 60 served and 10 refused.
"""
from starlette.datastructures import Headers
from backend.app import client_key
class _Req:
"""Enough of a Request for the key function: headers and a client host."""
def __init__(self, headers: dict, host: str = "10.0.0.9"):
self.headers = Headers(headers)
self.client = type("C", (), {"host": host})()
self.scope = {"type": "http", "client": (host, 0),
"headers": [(k.lower().encode(), v.encode())
for k, v in headers.items()]}
def test_cloudflare_header_wins():
# Cloudflare sets CF-Connecting-IP and overwrites any client-supplied
# value, so it is trustworthy in a way a parsed XFF chain is not.
assert client_key(_Req({"cf-connecting-ip": "203.0.113.7"})) == "203.0.113.7"
def test_forwarded_for_is_the_fallback_and_takes_the_first_entry():
# Left-most is the original client; everything after it is proxies.
assert client_key(
_Req({"x-forwarded-for": "203.0.113.7, 10.0.0.2"})) == "203.0.113.7"
def test_remote_address_is_the_last_resort():
assert client_key(_Req({}, host="10.0.0.9")) == "10.0.0.9"
def test_cloudflare_header_beats_forwarded_for():
key = client_key(_Req({"cf-connecting-ip": "203.0.113.7",
"x-forwarded-for": "198.51.100.1"}))
assert key == "203.0.113.7"
def test_two_callers_get_two_buckets():
# The whole point: one user exhausting their limit must not refuse another.
a = client_key(_Req({"cf-connecting-ip": "203.0.113.7"}))
b = client_key(_Req({"cf-connecting-ip": "203.0.113.8"}))
assert a != b
def test_whitespace_is_stripped():
# "a, b" split on comma leaves a leading space on every entry but the
# first; an unstripped key silently creates a second bucket per client.
assert client_key(_Req({"x-forwarded-for": " 203.0.113.7 ,10.0.0.2"})) \
== "203.0.113.7"
# ---------------------------------------------------------------------------
# The ceiling that header rotation cannot raise.
# ---------------------------------------------------------------------------
import pytest
from fastapi.testclient import TestClient
@pytest.fixture()
def api(monkeypatch):
from backend import app as app_module
from backend.config import settings
monkeypatch.setattr(settings, "global_rate_limit_per_minute", 5)
monkeypatch.setattr(app_module, "_global_window", None)
return TestClient(app_module.app, raise_server_exceptions=False)
def _ceiling_req(path: str, host: str):
"""Enough of a Request for exempt_from_ceiling: a path and a peer host."""
return type("R", (), {
"url": type("U", (), {"path": path})(),
"client": type("C", (), {"host": host})(),
})()
def _get(client, path="/api/flags", cf=None):
headers = {"cf-connecting-ip": cf} if cf else {}
return client.get(path, headers=headers)
def test_rotating_the_cloudflare_header_cannot_buy_unlimited_requests(api):
"""The attack the per-client keying opened up.
client_key trusts CF-Connecting-IP, and nothing in this process can tell an
edge-set header from an attacker-set one — that distinction can only be
made at Cloudflare, with Authenticated Origin Pulls or an origin firewall.
A caller reaching the origin directly can therefore mint a fresh
rate-limit bucket per request and evade per-client limits entirely.
Per-client fairness is still the right default; this is the backstop that
bounds what evading it can achieve. Without it, correct keying would be a
net regression against abuse compared with the shared bucket it replaced.
"""
codes = [_get(api, cf=f"203.0.113.{i}").status_code for i in range(8)]
assert codes.count(200) == 5
assert codes.count(429) == 3
def test_the_ceiling_says_which_limit_was_hit(api):
# Distinguishable from slowapi's per-client 429, or an operator reading
# logs cannot tell "one noisy client" from "the origin is saturated".
for i in range(5):
_get(api, cf=f"203.0.113.{i}")
refused = _get(api, cf="203.0.113.99")
assert refused.status_code == 429
assert "capacity" in refused.json()["detail"].lower()
assert refused.headers.get("retry-after")
def test_traffic_below_the_ceiling_is_untouched(api):
codes = [_get(api, cf=f"203.0.113.{i}").status_code for i in range(5)]
assert codes == [200] * 5
def test_the_container_healthcheck_is_exempt(api):
"""The healthcheck runs `curl http://localhost:80/api/data-info` inside the
container. If the ceiling could starve it, saturation would fail the
healthcheck, restart the container, and turn a load spike into an outage
loop — the ceiling has to protect the origin, not kill it.
"""
from backend.app import exempt_from_ceiling
assert exempt_from_ceiling(_ceiling_req("/api/data-info", "127.0.0.1"))
assert exempt_from_ceiling(_ceiling_req("/api/data-info", "::1"))
# Everyone else is counted.
assert not exempt_from_ceiling(_ceiling_req("/api/data-info", "10.0.0.9"))
def test_the_ceiling_ignores_non_api_paths():
# Sitemaps and robots.txt are served by this app too, and a crawler
# fetching them must not be refused because the API is busy.
from backend.app import exempt_from_ceiling
assert exempt_from_ceiling(_ceiling_req("/sitemap.xml", "10.0.0.9"))
assert exempt_from_ceiling(_ceiling_req("/robots.txt", "10.0.0.9"))
+16
View File
@@ -288,3 +288,19 @@ def test_outcodes_get_no_phase_variants(place_sitemaps):
# Nobody searches "primary schools in CM13"; the routes do not exist. # Nobody searches "primary schools in CM13"; the routes do not exist.
xml = place_sitemaps["outcodes-1.xml"] xml = place_sitemaps["outcodes-1.xml"]
assert "/primary" not in xml and "/secondary" not in xml assert "/primary" not in xml and "/secondary" not in xml
def test_authority_phase_variants_are_submitted_in_their_own_namespace(place_sitemaps):
"""302 of these were already in the sitemap, and every one 404'd.
The spec gives authorities a phase route; the plan built the bare
authority route and dropped it. Nothing noticed because the sitemap was
written from the registry, which was right, while the routes were written
by hand. This test fails if the URL ever leaves the sitemap; the e2e
journey fails if the route ever leaves the app.
"""
xml = place_sitemaps["places-1.xml"]
assert ("<loc>https://www.schoolcompare.co.uk"
"/schools/authority/essex/primary</loc>") in xml
# And never in the town namespace, which is a different set of schools.
assert "/schools/essex/primary" not in xml
+157
View File
@@ -0,0 +1,157 @@
"""Tests for school autosuggest (spec 2026-08-26)."""
from backend import data_loader
class _FakeDocs:
def __init__(self, hits, explode=False):
self._hits = hits
self._explode = explode
self.last_params = None
def search(self, params):
self.last_params = params
if self._explode:
raise RuntimeError("typesense is down")
return {"hits": [{"document": d} for d in self._hits]}
class _FakeClient:
def __init__(self, hits, explode=False):
self.docs = _FakeDocs(hits, explode)
self.collections = {"schools": type("C", (), {"documents": self.docs})()}
_HIT = {
"urn": 100010, "school_name": "Brecknock Primary School",
"local_authority": "Camden", "postcode": "NW1 1AA",
"phase": "Primary", "school_type": "Community school",
}
def _use(monkeypatch, client):
monkeypatch.setattr(data_loader, "_get_typesense_client", lambda: client)
def test_returns_the_fields_a_suggestion_needs(monkeypatch):
# Local authority is not decoration: there are many schools called
# "St Mary's", and a list without it cannot be chosen between.
_use(monkeypatch, _FakeClient([_HIT]))
out = data_loader.suggest_schools_typesense("breck")
assert out == [{
"urn": 100010, "school_name": "Brecknock Primary School",
"local_authority": "Camden", "postcode": "NW1 1AA",
"phase": "Primary", "school_type": "Community school",
}]
def test_a_missing_optional_field_becomes_an_empty_string(monkeypatch):
# phase and school_type are optional in the Typesense schema. A missing
# key must not KeyError in the keystroke path.
_use(monkeypatch, _FakeClient([{"urn": 1, "school_name": "X",
"local_authority": "Y", "postcode": "Z"}]))
out = data_loader.suggest_schools_typesense("x")
assert out[0]["phase"] == "" and out[0]["school_type"] == ""
def test_typesense_unavailable_gives_no_suggestions_rather_than_raising(monkeypatch):
_use(monkeypatch, None)
assert data_loader.suggest_schools_typesense("anything") == []
def test_a_typesense_error_gives_no_suggestions_rather_than_raising(monkeypatch):
_use(monkeypatch, _FakeClient([], explode=True))
assert data_loader.suggest_schools_typesense("anything") == []
def test_the_limit_is_passed_through_and_clamped(monkeypatch):
client = _FakeClient([])
_use(monkeypatch, client)
data_loader.suggest_schools_typesense("x", limit=500)
assert client.docs.last_params["per_page"] == 20
def _client(monkeypatch, rows, *, blow_up_dataframe=False):
from fastapi.testclient import TestClient
from backend import app as app_module
monkeypatch.setattr(app_module, "suggest_schools_typesense",
lambda q, limit=8: rows)
if blow_up_dataframe:
def _boom():
raise AssertionError("the suggest path must not load the DataFrame")
monkeypatch.setattr(app_module, "load_school_data", _boom)
monkeypatch.setattr(app_module, "load_latest_school_data", _boom)
return TestClient(app_module.app, raise_server_exceptions=False)
def test_the_endpoint_returns_suggestions(monkeypatch):
body = _client(monkeypatch, [_HIT]).get("/api/suggest?q=breck").json()
assert body["suggestions"][0]["school_name"] == "Brecknock Primary School"
def test_the_endpoint_never_touches_the_dataframe(monkeypatch):
"""The whole reason this is not a mode of /api/schools.
That endpoint filters and sorts 25,000 rows of pandas per query, holding
the GIL. Per keystroke, that is the cost this endpoint exists to avoid.
"""
res = _client(monkeypatch, [_HIT], blow_up_dataframe=True).get("/api/suggest?q=breck")
assert res.status_code == 200
assert res.json()["suggestions"]
def test_a_one_character_query_returns_nothing_and_does_not_error(monkeypatch):
# The keystroke path never errors on ordinary input.
res = _client(monkeypatch, [_HIT]).get("/api/suggest?q=b")
assert res.status_code == 200
assert res.json() == {"suggestions": []}
def test_a_blank_query_returns_nothing_and_does_not_error(monkeypatch):
res = _client(monkeypatch, [_HIT]).get("/api/suggest?q=")
assert res.status_code == 200
assert res.json() == {"suggestions": []}
def test_typesense_down_is_an_empty_list_not_a_500(monkeypatch):
res = _client(monkeypatch, []).get("/api/suggest?q=breck")
assert res.status_code == 200
assert res.json() == {"suggestions": []}
def test_the_response_is_cacheable(monkeypatch):
# Prefix queries repeat enormously across users, and school names change
# once a year. Without this the endpoint pays full price every keystroke.
res = _client(monkeypatch, [_HIT]).get("/api/suggest?q=breck")
assert "s-maxage" in res.headers.get("cache-control", "")
assert res.headers.get("etag")
def test_a_malformed_urn_does_not_raise(monkeypatch):
"""The docstring promises "never raises"; the parsing loop sat outside the
try, so int(None) or int("abc") would have turned a keystroke into a 500.
Typesense declares urn as int32, so this should be unreachable — but the
contract is what the caller relies on, and a search index is a separate
system that can be reindexed by something other than this code.
"""
_use(monkeypatch, _FakeClient([{"urn": None, "school_name": "X",
"local_authority": "Y", "postcode": "Z"}]))
assert data_loader.suggest_schools_typesense("x") == []
def test_a_malformed_row_does_not_discard_the_good_ones(monkeypatch):
# One bad document must not blank the whole dropdown.
_use(monkeypatch, _FakeClient([
{"urn": "not-a-number", "school_name": "Bad", "local_authority": "Y",
"postcode": "Z"},
_HIT,
]))
out = data_loader.suggest_schools_typesense("x")
assert [r["urn"] for r in out] == [100010]
def test_a_hit_with_no_document_does_not_raise(monkeypatch):
_use(monkeypatch, _FakeClient([{}]))
assert data_loader.suggest_schools_typesense("x") == []
+9 -2
View File
@@ -132,16 +132,23 @@ def test_one_query_per_table_and_latest_row_per_urn():
"FactPupilCharacteristics": [], "FactPupilCharacteristics": [],
"FactDeprivation": [], "FactDeprivation": [],
"FactFinance": [], "FactFinance": [],
"FactKs4Destinations": [],
"FactKs5Destinations": [],
} }
session = _FakeSession(rows) session = _FakeSession(rows)
out = get_supplementary_data_batch(session, [1, 2]) out = get_supplementary_data_batch(session, [1, 2])
# Exactly one query per table — six total, regardless of two URNs. # Exactly one query per table — eight total, regardless of two URNs.
assert sorted(session.queries) == [ assert sorted(session.queries) == [
"FactAdmissionDistance", "FactAdmissions", "FactDeprivation", "FactAdmissionDistance", "FactAdmissions", "FactDeprivation",
"FactFinance", "FactOfstedInspection", "FactPupilCharacteristics", "FactFinance", "FactKs4Destinations", "FactKs5Destinations",
"FactOfstedInspection", "FactPupilCharacteristics",
] ]
# A school with no destination rows gets null, not an empty shell — the
# frontend renders the section from the block's presence.
assert out[1]["destinations"] is None
# Latest Ofsted kept per URN # Latest Ofsted kept per URN
assert out[1]["ofsted"]["overall_effectiveness"] == 2 assert out[1]["ofsted"]["overall_effectiveness"] == 2
assert out[2]["ofsted"]["overall_effectiveness"] == 1 assert out[2]["ofsted"]["overall_effectiveness"] == 1
+46
View File
@@ -23,6 +23,52 @@ Key files:
- `backend/data_loader.py` - Data queries, geocoding, legacy DataFrame compatibility - `backend/data_loader.py` - Data queries, geocoding, legacy DataFrame compatibility
- `backend/schemas.py` - Column mappings, metric definitions, LA code mappings - `backend/schemas.py` - Column mappings, metric definitions, LA code mappings
### Content / CMS (Payload)
Payload CMS runs **inside** the Next.js app — one image, one container, no
separate service. It powers `/blog`; `/about` is a plain coded page.
- **Admin panel:** `/admin`. The only authenticated surface on the site.
`noindex` via both `robots.txt` and `X-Robots-Tag`.
- **CMS API:** `/cms-api`, **not** `/api`. `/api/*` is a catch-all proxy to
FastAPI (`app/(frontend)/api/[...path]`) which would silently swallow every
admin call and forward it to the backend. Mount points are defined once in
`lib/payloadRoutes.ts`.
- **Database:** the existing Postgres, in its own `payload` schema, so no
pipeline operation on `public` — including
`scripts/migrate_csv_to_db.py --drop` — can reach blog content.
- **Uploads:** the `payload_media` Docker volume at `/app/media`. Not
reproducible from the pipeline; must be backed up.
- **New env vars:** `DATABASE_URL` and `PAYLOAD_SECRET` on the frontend service.
Staging must use a different `PAYLOAD_SECRET` from production.
- Publishing workflow and house style: `nextjs-app/docs/PUBLISHING.md`.
- **Admin field components resolve through a generated import map**
(`app/(payload)/admin/importMap.js`). Payload hands the client a *path* per
field and looks it up there; a missing entry renders no field and reports no
error, while `required` still blocks the save. After adding or changing any
field, editor or lexical feature, run `npm run generate:importmap` in
`nextjs-app/` and commit the result.
### Two route groups
`nextjs-app/app/` has no root `layout.tsx`. It cannot: Payload's admin panel
ships its own root layout rendering `<html>`/`<body>`, and Next permits
multiple root layouts only when no `app/layout.tsx` exists.
- `app/(frontend)/` — the site. Its `layout.tsx` is the site's root layout.
- `app/(payload)/` — the admin panel and `/cms-api`.
Route groups are invisible to routing, so every public URL is unchanged.
**The metadata file conventions stay at the `app/` root** — `robots.ts`,
`opengraph-image.tsx`, `icon.png`, `apple-icon.png`. Inside a route group Next
treats them as segment-scoped: it renames `/icon.png` to `/icon-<hash>.png` and
drops `/robots.txt` entirely. Route handlers are unaffected.
The build must succeed with `DATABASE_URL` unset, because CI builds it that
way. Never call `getCachedPayload()` at module scope, and never add
`generateStaticParams` to a DB-backed route.
### Frontend (Vanilla JS) ### Frontend (Vanilla JS)
- Single-page application with hash-based routing - Single-page application with hash-based routing
- Chart.js for data visualization - Chart.js for data visualization
+48 -2
View File
@@ -16,7 +16,16 @@
# ADMIN_API_KEY — Backend admin API key # ADMIN_API_KEY — Backend admin API key
# TYPESENSE_API_KEY — Typesense admin API key # TYPESENSE_API_KEY — Typesense admin API key
# TYPESENSE_SEARCH_KEY — Typesense search-only key (exposed to frontend) # TYPESENSE_SEARCH_KEY — Typesense search-only key (exposed to frontend)
# AIRFLOW_ADMIN_USER — Airflow admin username (password auto-generated, see api-server logs) # UNLEASH_URL — http://<unleash-ip>:4242/api (empty = all flags off)
# UNLEASH_API_TOKEN — Unleash *client* token, environment: development
# PAYLOAD_SECRET — Payload CMS encryption secret. REQUIRED: long and
# random, and DIFFERENT from production's. Sharing
# it would let a staging session authenticate
# against production.
# AIRFLOW_ADMIN_USER — Airflow admin username (default: admin)
# AIRFLOW_ADMIN_PASSWORD — Airflow admin password. REQUIRED: the api-server
# refuses to start without it, rather than falling
# back to a generated one that changes on restart.
# STAGING_DB_IP — macvlan IP for staging Postgres (default 10.0.1.190) # STAGING_DB_IP — macvlan IP for staging Postgres (default 10.0.1.190)
# STAGING_FRONTEND_IP — macvlan IP for staging frontend (default 10.0.1.151) # STAGING_FRONTEND_IP — macvlan IP for staging frontend (default 10.0.1.151)
@@ -55,6 +64,12 @@ services:
ADMIN_API_KEY: ${ADMIN_API_KEY:-changeme} ADMIN_API_KEY: ${ADMIN_API_KEY:-changeme}
TYPESENSE_URL: http://typesense:8108 TYPESENSE_URL: http://typesense:8108
TYPESENSE_API_KEY: ${TYPESENSE_API_KEY:-changeme} TYPESENSE_API_KEY: ${TYPESENSE_API_KEY:-changeme}
# Unset means every feature flag is False — the correct dark state for an
# environment with no Unleash, not a failure.
UNLEASH_URL: ${UNLEASH_URL:-}
UNLEASH_API_TOKEN: ${UNLEASH_API_TOKEN:-}
volumes:
- unleash_cache:/app/.unleash
depends_on: depends_on:
sc_database: sc_database:
condition: service_healthy condition: service_healthy
@@ -78,9 +93,20 @@ services:
- FASTAPI_URL=http://backend:80/api - FASTAPI_URL=http://backend:80/api
- TYPESENSE_URL=http://typesense:8108 - TYPESENSE_URL=http://typesense:8108
- TYPESENSE_API_KEY=${TYPESENSE_SEARCH_KEY:-changeme} - TYPESENSE_API_KEY=${TYPESENSE_SEARCH_KEY:-changeme}
# Payload CMS runs inside this container, in the `payload` schema of the
# staging database. Staging has its own stack, its own Postgres and its
# own admin account — never production's.
- DATABASE_URL=postgresql://${DB_USERNAME}:${DB_PASSWORD}@sc_database:5432/${DB_DATABASE_NAME}
- PAYLOAD_SECRET=${PAYLOAD_SECRET:?set PAYLOAD_SECRET in the staging Portainer stack environment}
volumes:
# Portainer prefixes volume names with the stack name, so this is
# automatically isolated from production's media.
- payload_media:/app/media
depends_on: depends_on:
backend: backend:
condition: service_healthy condition: service_healthy
sc_database:
condition: service_healthy
networks: networks:
backend: {} backend: {}
macvlan: macvlan:
@@ -116,7 +142,23 @@ services:
airflow-api-server: airflow-api-server:
image: privaterepo.sitaru.org/tudor/school_compare-pipeline:staging image: privaterepo.sitaru.org/tudor/school_compare-pipeline:staging
container_name: sc_staging_airflow_api container_name: sc_staging_airflow_api
command: airflow api-server --port 8080 # The simple auth manager generates a random password on first start and
# writes it to a file, so every container restart invalidates the last one.
# Writing the file ourselves from an environment variable makes the login
# deterministic. Airflow does not generate anything when the file exists.
#
# Built with python rather than echo/printf so a password containing quotes,
# backslashes or spaces is escaped correctly by json.dumps. An unset
# AIRFLOW_ADMIN_PASSWORD raises KeyError and the container exits: falling
# back to a generated password would silently undo the point of this.
command:
- bash
- -c
- |
set -euo pipefail
mkdir -p /opt/airflow
python -c "import json, os, pathlib; pathlib.Path('/opt/airflow/simple_auth_manager_passwords.json').write_text(json.dumps({os.environ.get('AIRFLOW_ADMIN_USER', 'admin'): os.environ['AIRFLOW_ADMIN_PASSWORD']}))"
exec airflow api-server --port 8080
ports: ports:
- "8081:8080" - "8081:8080"
environment: environment:
@@ -128,6 +170,8 @@ services:
AIRFLOW__API_AUTH__JWT_SECRET: "school-compare-staging-airflow-jwt-secret-key-long-enough-for-sha512" AIRFLOW__API_AUTH__JWT_SECRET: "school-compare-staging-airflow-jwt-secret-key-long-enough-for-sha512"
AIRFLOW__API_AUTH__JWT_ISSUER: airflow AIRFLOW__API_AUTH__JWT_ISSUER: airflow
AIRFLOW__CORE__SIMPLE_AUTH_MANAGER_USERS: "${AIRFLOW_ADMIN_USER:-admin}:admin" AIRFLOW__CORE__SIMPLE_AUTH_MANAGER_USERS: "${AIRFLOW_ADMIN_USER:-admin}:admin"
AIRFLOW__CORE__SIMPLE_AUTH_MANAGER_PASSWORDS_FILE: /opt/airflow/simple_auth_manager_passwords.json
AIRFLOW_ADMIN_PASSWORD: ${AIRFLOW_ADMIN_PASSWORD:?set AIRFLOW_ADMIN_PASSWORD in the Portainer stack environment}
AIRFLOW__LOGGING__BASE_LOG_FOLDER: /opt/airflow/logs AIRFLOW__LOGGING__BASE_LOG_FOLDER: /opt/airflow/logs
PG_HOST: sc_database PG_HOST: sc_database
PG_PORT: "5432" PG_PORT: "5432"
@@ -212,3 +256,5 @@ volumes:
postgres_data: postgres_data:
typesense_data: typesense_data:
airflow_logs: airflow_logs:
unleash_cache:
payload_media:
+73
View File
@@ -0,0 +1,73 @@
# Portainer Stack Definition for School Compare — UNLEASH (feature flags)
#
# Deploy as a *separate* Portainer stack ("schoolcompare-unleash"), alongside
# the production and staging stacks. It deliberately belongs to neither: a
# staging redeploy must not be able to disturb production's flag state, and a
# production redeploy must not disturb staging's.
#
# One instance serves both environments. Open-source Unleash ships with
# `development` and `production` environments and environment-scoped client
# tokens, so the same flag holds independent state in each — which is what
# lets a feature be on in staging, where the E2E journeys exercise it, while
# production stays dark.
#
# Portainer environment variables (set in Portainer UI -> Stack -> Environment):
# UNLEASH_DB_PASSWORD — PostgreSQL password for the Unleash database
# UNLEASH_ADMIN_PASSWORD — initial admin password for the Unleash UI
# UNLEASH_IP — macvlan IP for the Unleash server (default 10.0.1.152)
services:
# ── PostgreSQL (Unleash's own; nothing else uses it) ──────────────────
unleash_db:
container_name: sc_unleash_postgres
image: postgres:16-alpine
environment:
POSTGRES_USER: unleash
POSTGRES_PASSWORD: ${UNLEASH_DB_PASSWORD}
POSTGRES_DB: unleash
volumes:
- unleash_postgres_data:/var/lib/postgresql/data
networks:
- unleash
healthcheck:
test: ["CMD-SHELL", "pg_isready -U unleash"]
interval: 10s
timeout: 5s
retries: 5
start_period: 10s
restart: unless-stopped
# ── Unleash server (UI + client API on 4242) ──────────────────────────
unleash:
container_name: sc_unleash
image: unleashorg/unleash-server:6
environment:
DATABASE_URL: postgres://unleash:${UNLEASH_DB_PASSWORD}@unleash_db:5432/unleash
DATABASE_SSL: "false"
INIT_ADMIN_API_TOKENS: ""
UNLEASH_DEFAULT_ADMIN_PASSWORD: ${UNLEASH_ADMIN_PASSWORD}
depends_on:
unleash_db:
condition: service_healthy
networks:
unleash: {}
macvlan:
ipv4_address: ${UNLEASH_IP:-10.0.1.152}
healthcheck:
test: ["CMD-SHELL", "wget -qO- http://localhost:4242/health || exit 1"]
interval: 30s
timeout: 10s
retries: 3
start_period: 30s
restart: unless-stopped
networks:
unleash:
driver: bridge
macvlan:
external:
name: macvlan
volumes:
unleash_postgres_data:
+48 -2
View File
@@ -7,7 +7,15 @@
# ADMIN_API_KEY — Backend admin API key # ADMIN_API_KEY — Backend admin API key
# TYPESENSE_API_KEY — Typesense admin API key # TYPESENSE_API_KEY — Typesense admin API key
# TYPESENSE_SEARCH_KEY — Typesense search-only key (exposed to frontend) # TYPESENSE_SEARCH_KEY — Typesense search-only key (exposed to frontend)
# AIRFLOW_ADMIN_USER — Airflow admin username (password auto-generated, see api-server logs) # UNLEASH_URL — http://<unleash-ip>:4242/api (empty = all flags off)
# UNLEASH_API_TOKEN — Unleash *client* token, environment: production
# PAYLOAD_SECRET — Payload CMS encryption secret. REQUIRED: long and
# random. Changing it invalidates every admin
# session. Staging MUST use a different value.
# AIRFLOW_ADMIN_USER — Airflow admin username (default: admin)
# AIRFLOW_ADMIN_PASSWORD — Airflow admin password. REQUIRED: the api-server
# refuses to start without it, rather than falling
# back to a generated one that changes on restart.
services: services:
@@ -44,6 +52,12 @@ services:
ADMIN_API_KEY: ${ADMIN_API_KEY:-changeme} ADMIN_API_KEY: ${ADMIN_API_KEY:-changeme}
TYPESENSE_URL: http://typesense:8108 TYPESENSE_URL: http://typesense:8108
TYPESENSE_API_KEY: ${TYPESENSE_API_KEY:-changeme} TYPESENSE_API_KEY: ${TYPESENSE_API_KEY:-changeme}
# Unset means every feature flag is False — the correct dark state for an
# environment with no Unleash, not a failure.
UNLEASH_URL: ${UNLEASH_URL:-}
UNLEASH_API_TOKEN: ${UNLEASH_API_TOKEN:-}
volumes:
- unleash_cache:/app/.unleash
depends_on: depends_on:
sc_database: sc_database:
condition: service_healthy condition: service_healthy
@@ -67,9 +81,21 @@ services:
- FASTAPI_URL=http://backend:80/api - FASTAPI_URL=http://backend:80/api
- TYPESENSE_URL=http://typesense:8108 - TYPESENSE_URL=http://typesense:8108
- TYPESENSE_API_KEY=${TYPESENSE_SEARCH_KEY:-changeme} - TYPESENSE_API_KEY=${TYPESENSE_SEARCH_KEY:-changeme}
# Payload CMS runs inside this container. It reaches Postgres over the
# `backend` network and keeps its tables in the `payload` schema, so no
# pipeline operation on `public` can touch blog content.
- DATABASE_URL=postgresql://${DB_USERNAME}:${DB_PASSWORD}@sc_database:5432/${DB_DATABASE_NAME}
# Same :? form as AIRFLOW_ADMIN_PASSWORD: refuse to start rather than
# boot with an empty secret and silently accept forged sessions.
- PAYLOAD_SECRET=${PAYLOAD_SECRET:?set PAYLOAD_SECRET in the Portainer stack environment}
volumes:
# Blog images. Not reproducible from the pipeline — must be backed up.
- payload_media:/app/media
depends_on: depends_on:
backend: backend:
condition: service_healthy condition: service_healthy
sc_database:
condition: service_healthy
networks: networks:
backend: {} backend: {}
macvlan: macvlan:
@@ -105,7 +131,23 @@ services:
airflow-api-server: airflow-api-server:
image: privaterepo.sitaru.org/tudor/school_compare-pipeline:prod image: privaterepo.sitaru.org/tudor/school_compare-pipeline:prod
container_name: schoolcompare_airflow_api container_name: schoolcompare_airflow_api
command: airflow api-server --port 8080 # The simple auth manager generates a random password on first start and
# writes it to a file, so every container restart invalidates the last one.
# Writing the file ourselves from an environment variable makes the login
# deterministic. Airflow does not generate anything when the file exists.
#
# Built with python rather than echo/printf so a password containing quotes,
# backslashes or spaces is escaped correctly by json.dumps. An unset
# AIRFLOW_ADMIN_PASSWORD raises KeyError and the container exits: falling
# back to a generated password would silently undo the point of this.
command:
- bash
- -c
- |
set -euo pipefail
mkdir -p /opt/airflow
python -c "import json, os, pathlib; pathlib.Path('/opt/airflow/simple_auth_manager_passwords.json').write_text(json.dumps({os.environ.get('AIRFLOW_ADMIN_USER', 'admin'): os.environ['AIRFLOW_ADMIN_PASSWORD']}))"
exec airflow api-server --port 8080
ports: ports:
- "8080:8080" - "8080:8080"
environment: environment:
@@ -117,6 +159,8 @@ services:
AIRFLOW__API_AUTH__JWT_SECRET: "school-compare-airflow-jwt-secret-key-long-enough-for-sha512" AIRFLOW__API_AUTH__JWT_SECRET: "school-compare-airflow-jwt-secret-key-long-enough-for-sha512"
AIRFLOW__API_AUTH__JWT_ISSUER: airflow AIRFLOW__API_AUTH__JWT_ISSUER: airflow
AIRFLOW__CORE__SIMPLE_AUTH_MANAGER_USERS: "${AIRFLOW_ADMIN_USER:-admin}:admin" AIRFLOW__CORE__SIMPLE_AUTH_MANAGER_USERS: "${AIRFLOW_ADMIN_USER:-admin}:admin"
AIRFLOW__CORE__SIMPLE_AUTH_MANAGER_PASSWORDS_FILE: /opt/airflow/simple_auth_manager_passwords.json
AIRFLOW_ADMIN_PASSWORD: ${AIRFLOW_ADMIN_PASSWORD:?set AIRFLOW_ADMIN_PASSWORD in the Portainer stack environment}
AIRFLOW__LOGGING__BASE_LOG_FOLDER: /opt/airflow/logs AIRFLOW__LOGGING__BASE_LOG_FOLDER: /opt/airflow/logs
PG_HOST: sc_database PG_HOST: sc_database
PG_PORT: "5432" PG_PORT: "5432"
@@ -201,3 +245,5 @@ volumes:
postgres_data: postgres_data:
typesense_data: typesense_data:
airflow_logs: airflow_logs:
unleash_cache:
payload_media:
+23 -1
View File
@@ -36,6 +36,10 @@ services:
ADMIN_API_KEY: ${ADMIN_API_KEY:-changeme} ADMIN_API_KEY: ${ADMIN_API_KEY:-changeme}
TYPESENSE_URL: http://typesense:8108 TYPESENSE_URL: http://typesense:8108
TYPESENSE_API_KEY: ${TYPESENSE_API_KEY:-changeme} TYPESENSE_API_KEY: ${TYPESENSE_API_KEY:-changeme}
# Unset means every feature flag is False — the correct dark state for an
# environment with no Unleash, not a failure.
UNLEASH_URL: ${UNLEASH_URL:-}
UNLEASH_API_TOKEN: ${UNLEASH_API_TOKEN:-}
volumes: volumes:
- ./data:/app/data:ro - ./data:/app/data:ro
depends_on: depends_on:
@@ -101,7 +105,23 @@ services:
airflow-api-server: airflow-api-server:
image: privaterepo.sitaru.org/tudor/school_compare-pipeline:latest image: privaterepo.sitaru.org/tudor/school_compare-pipeline:latest
container_name: schoolcompare_airflow_api container_name: schoolcompare_airflow_api
command: airflow api-server --port 8080 # The simple auth manager generates a random password on first start and
# writes it to a file, so every container restart invalidates the last one.
# Writing the file ourselves from an environment variable makes the login
# deterministic. Airflow does not generate anything when the file exists.
#
# Built with python rather than echo/printf so a password containing quotes,
# backslashes or spaces is escaped correctly by json.dumps. An unset
# AIRFLOW_ADMIN_PASSWORD raises KeyError and the container exits: falling
# back to a generated password would silently undo the point of this.
command:
- bash
- -c
- |
set -euo pipefail
mkdir -p /opt/airflow
python -c "import json, os, pathlib; pathlib.Path('/opt/airflow/simple_auth_manager_passwords.json').write_text(json.dumps({os.environ.get('AIRFLOW_ADMIN_USER', 'admin'): os.environ['AIRFLOW_ADMIN_PASSWORD']}))"
exec airflow api-server --port 8080
ports: ports:
- "8080:8080" - "8080:8080"
environment: &airflow-env environment: &airflow-env
@@ -113,6 +133,8 @@ services:
AIRFLOW__API_AUTH__JWT_SECRET: "school-compare-airflow-jwt-secret-key-long-enough-for-sha512" AIRFLOW__API_AUTH__JWT_SECRET: "school-compare-airflow-jwt-secret-key-long-enough-for-sha512"
AIRFLOW__API_AUTH__JWT_ISSUER: airflow AIRFLOW__API_AUTH__JWT_ISSUER: airflow
AIRFLOW__CORE__SIMPLE_AUTH_MANAGER_USERS: "admin:admin" AIRFLOW__CORE__SIMPLE_AUTH_MANAGER_USERS: "admin:admin"
AIRFLOW__CORE__SIMPLE_AUTH_MANAGER_PASSWORDS_FILE: /opt/airflow/simple_auth_manager_passwords.json
AIRFLOW_ADMIN_PASSWORD: ${AIRFLOW_ADMIN_PASSWORD:-admin}
PG_HOST: db PG_HOST: db
PG_PORT: "5432" PG_PORT: "5432"
PG_USER: schoolcompare PG_USER: schoolcompare
+106
View File
@@ -98,6 +98,12 @@ fail the E2E gate. That's the point: staging absorbs the risk.
pr-checks status checks (frontend, backend, builds, ai-review) to pass. pr-checks status checks (frontend, backend, builds, ai-review) to pass.
5. **Bootstrap staging data via Airflow** (no prod dump — staging populates 5. **Bootstrap staging data via Airflow** (no prod dump — staging populates
itself from source, exercising the pipeline image end-to-end): itself from source, exercising the pipeline image end-to-end):
- Set `AIRFLOW_ADMIN_PASSWORD` in the stack environment first. The
api-server refuses to start without it. Airflow's simple auth manager
otherwise generates a password on first start and writes it to a file, so
the login changes every time the container restarts; the stack writes that
file itself from this variable instead. `AIRFLOW_ADMIN_USER` defaults to
`admin`.
- Open the staging Airflow UI (`http://<host>:8081`) and trigger, in order: - Open the staging Airflow UI (`http://<host>:8081`) and trigger, in order:
`school_data_daily`, `school_data_monthly_ofsted`, then the manual-schedule `school_data_daily`, `school_data_monthly_ofsted`, then the manual-schedule
`school_data_annual_ees` and `school_data_annual_idaci`. `school_data_annual_ees` and `school_data_annual_idaci`.
@@ -150,3 +156,103 @@ token Gitea Actions provides automatically (`secrets.GITEA_TOKEN` — no setup
needed), and fails the check only when a finding is rated needed), and fails the check only when a finding is rated
**severe** (would break prod, leak data, or corrupt data). Minor findings are **severe** (would break prod, leak data, or corrupt data). Minor findings are
informational and never block a merge. informational and never block a merge.
## Rate limiting, and the Cloudflare gap
Two independent limits protect the API:
- **Per client**, via slowapi, keyed on `CF-Connecting-IP` (falling back to
`X-Forwarded-For`, then the peer address). 60/minute by default;
`/api/suggest` gets 120/minute because typing is bursty.
- **Globally**, via `GlobalRateLimitMiddleware`: a fixed 60-second window over
all `/api/` traffic, `GLOBAL_RATE_LIMIT_PER_MINUTE` (default 3000),
independent of any client identity. Requests from `127.0.0.1` are exempt so
the container healthcheck cannot be starved into a restart loop.
### Open: the origin must only accept Cloudflare
`CF-Connecting-IP` is only meaningful for requests that actually reached the
origin through Cloudflare, and **the application cannot verify that they did**.
Anything able to reach the origin directly can set that header freely and, by
rotating it, mint a fresh rate-limit bucket per request — defeating per-client
limits on every endpoint.
The global ceiling bounds the damage to total origin capacity. It does not fix
the underlying gap, and nothing in the code can. Closing it needs one of:
- **Authenticated Origin Pulls** — Cloudflare presents a client certificate the
origin requires, so non-Cloudflare traffic is refused at TLS.
- **An origin firewall** restricted to Cloudflare's published IP ranges.
Until one is in place, treat per-client limits as protection against accidents
and ordinary load, not against a determined caller.
## Feature flags (Unleash)
Flag state lives in a self-hosted Unleash instance, deployed as its own
Portainer stack from `docker-compose.portainer.unleash.yml`. It is separate
from the application stacks on purpose — redeploying staging must not be able
to disturb production's flags.
The flags themselves are declared in `backend/flags.py`. Unleash holds the
state; the registry holds the list. A flag in the UI that is not in the
registry is orphaned and nothing reads it.
### First-time setup
1. Deploy the stack in Portainer. Set `UNLEASH_DB_PASSWORD`,
`UNLEASH_ADMIN_PASSWORD` and (optionally) `UNLEASH_IP`.
2. Log in to the UI at `http://<UNLEASH_IP>:4242` as `admin`.
3. Create one **client** API token per environment:
- `schoolcompare-staging`, environment **development**
- `schoolcompare-prod`, environment **production**
Client tokens, not admin tokens — the backend only reads.
4. Put each token in the matching Portainer stack's `UNLEASH_API_TOKEN`
variable, and set `UNLEASH_URL` to `http://<UNLEASH_IP>:4242/api`.
5. Redeploy the application stacks.
### Adding a flag to Unleash
**Unleash does not create flags by itself.** The SDK reads definitions from the
server and never registers anything, and metrics for a flag the server has
never heard of are discarded. So a flag declared in `backend/flags.py` will be
evaluated on every request, stay `False` forever, and never appear in the UI
until someone creates it there by hand.
For each flag in the registry, create one in Unleash with:
- **Name** — character for character what `backend/flags.py` declares.
snake_case, no hyphens or spaces. A typo produces a flag that looks correct
in the UI and is read by nothing.
- **Type** — Release. No strategies, constraints or variants: these are plain
on/off switches, by design.
### Turning a feature on
Toggle the flag in the environment matching the stack you mean: **development**
for staging, **production** for prod. The token in each stack is scoped to one
environment, so toggling the other one has no visible effect.
The SDK refreshes every 15 seconds, so the API reflects the change almost at
once; the pages follow on their own schedule, below.
A flip reaches school pages within about five minutes and place pages within
the hour. Next's ISR does the propagating — it revalidates a route at the
*lowest* `revalidate` among that route's fetches, which is 300s for
`/school/[slug]` and 3600s for the place pages. There is no webhook, and
adding one would only be worth it if flips ever needed to be instant.
### When Unleash is unreachable
Every flag evaluates to `False` and the site serves as though nothing were
switched on. That is deliberate — an unfinished feature staying hidden is the
safe direction — but it means a *released* feature disappears if a backend
container cold-starts with an empty cache while Unleash is down. The SDK's
disk cache is on a named volume so restarts keep last-known state, and flags
are removed from the code within 90 days (enforced by a test), which bounds
how long any feature is exposed to this.
If `UNLEASH_URL` is unset, every flag is `False` and no connection is
attempted. That is the correct behaviour for local development and CI, and it
means the test suites need no flag server.
File diff suppressed because it is too large. Load diff
File diff suppressed because it is too large. Load diff
File diff suppressed because it is too large. Load diff
File diff suppressed because it is too large. Load diff
@@ -0,0 +1,294 @@
# Feature Flags — Design
**Date:** 2026-08-23
**Status:** approved for planning
**First consumer:** the last-distance-offered feature (`admission_distance`)
## Goal
Let work merge to `main` and deploy to production without becoming visible,
so that releasing a feature stops being the same event as deploying it.
The site has no way to do this today. A feature is either on `main` and live,
or it is on a branch. That forces long-lived branches for anything not ready,
and it makes every promotion to production an all-or-nothing decision about
everything queued behind it.
This is a **ship-dark** capability, not a kill switch. Flags are expected to
flip on the order of once a month, by a person, deliberately. Nothing here is
designed for flipping something off in seconds under pressure, and nothing
here does percentage rollouts, user targeting or A/B tests — the site has no
user identity to target.
## Decision: Unleash
Flag state is held in a self-hosted [Unleash](https://www.getunleash.io/)
instance (Apache-2.0), not in the repository.
A lighter option was considered and rejected by the project owner: a typed
registry in each runtime with environment-variable overrides set in the
Portainer stack files, which would have needed no new container and kept flag
state in git. The argument for Unleash is that it provides a UI and an audit
log without a deploy, and that flags are expected to become an ongoing
operational tool rather than an occasional one.
Two consequences follow from choosing a service, and this design exists mostly
to handle them:
1. **Flag state lives outside the repository.** `main` is no longer the whole
truth about what is switched on. The registry in §2 exists to bound that.
2. **A flag can change without a deploy**, so nothing else clears the caches
that a deploy would have cleared. §4 establishes how long a flip takes to
become visible, and why that is short enough to need no extra mechanism.
Also considered: Flagsmith (heavier — Django, Postgres and Redis), GrowthBook
(requires MongoDB), and Flipt v2 (the closest conceptual fit, git-native, but
now under the Fair Core Licence — source-available, not OSI open source).
## 1. Topology
A third Portainer stack, `docker-compose.portainer.unleash.yml`, holding
`unleashorg/unleash-server` and its own PostgreSQL 16. It is on the macvlan so
both application stacks can reach it, and it belongs to neither of them — a
staging redeploy must not be able to disturb production's flag state, and vice
versa.
One instance serves both environments. Open-source Unleash ships with
`development` and `production` environments and environment-scoped client
tokens, so the same flag holds independent state in each: staging's FastAPI
carries a `development` token, production's carries a `production` one.
That property is what makes ship-dark testable. A feature can be **on in
staging and off in production** for as long as it takes, which means the `e2e/`
journeys exercise it against staging while production stays unchanged.
## 2. The registry
Unleash supplies flag *state* and the toggle UI. It does not supply the list of
flags. `backend/flags.py` declares every flag the code knows about:
```python
@dataclass(frozen=True)
class Flag:
name: str # identical in the registry, in Unleash, and in JSON
description: str # one line: what turning this on reveals
added: date # for the staleness test in §8
```
**Every flag defaults to `False`.** There is no per-flag default field, because
a flag that defaults on is not a ship-dark flag — it is a kill switch, and this
design does not offer one. A single unconditional default also means the
fallback path has no branching to get wrong.
Three reasons the registry is not optional:
- The Unleash SDK evaluates an unknown flag to `False`. Without a registry that
is an *undeclared* false — indistinguishable from a typo in a flag name.
- `/api/flags` needs a key set to return when Unleash is unreachable. It cannot
enumerate flags it has never heard of.
- A flag present in the Unleash UI but absent from the registry is orphaned,
and should be visibly so rather than quietly authoritative.
**Naming.** One string, used unchanged as the registry key, the Unleash flag
name, and the JSON key in `/api/flags`. It is snake_case, matching the API's
existing convention (`admission_distance`, `rwm_expected_pct`) and the mirrored
types in `nextjs-app/lib/types.ts`. No case transformation anywhere, so there
is no mapping layer to get wrong.
## 3. Read paths
### Backend
`backend/flags.py` wraps `UnleashClient` behind `is_enabled(name: str) -> bool`.
Fail-closed is the default rather than something added: the Python SDK
evaluates every flag to `False` until it has synchronised with the server. An
unfinished feature therefore stays hidden when Unleash is unreachable, which is
the correct direction for ship-dark.
The SDK's fcache directory is mounted on a named volume so a container restart
during an Unleash outage keeps last-known state rather than reverting a
released feature to dark. The registry default remains `False`, so the worst
case is a feature disappearing, never one appearing.
### Frontend
`nextjs-app/lib/flags.ts` exposes `getFlags(): Promise<Flags>`, a single
server-side fetch of `/api/flags` returning a typed record. Server components
only — no flag value reaches the browser bundle, and `package.json` gains no
Unleash dependency. The Unleash client library stays entirely inside the
service that already owns every other piece of data the frontend renders.
The cost, named plainly: a purely front-end flag must still be declared in a
Python file. It is a flat data edit rather than programming, and the return is
one list, so nobody has to ask which service knows about a given flag.
### `/api/flags` must not be publicly reachable
`nextjs-app/app/api/[...path]/route.ts` proxies **everything** under `/api/` to
FastAPI. Left alone, `https://www.schoolcompare.co.uk/api/flags` would return
`{"admission_distance": false, ...}` — publishing the name and state of every
unreleased feature, which defeats the purpose of shipping dark.
The proxy therefore gains a denylist, and `flags` is on it: a request for a
denied path returns 404 rather than being forwarded. Next's own `getFlags()` is
unaffected because it calls `FASTAPI_URL` directly across the Docker network
and never transits the public proxy.
This is a general hole rather than a flags-specific one — the proxy will
forward any future internal endpoint too — so the denylist is written as a
named constant with a comment saying what belongs on it.
## 4. Propagation
**Time-based revalidation is sufficient. There is no webhook.**
An earlier draft of this section specified two Unleash webhooks and a
`revalidateTag('flags')` purge, on the premise that pages cache for seven days.
That premise was wrong, and checking it removed the most complex part of the
design.
Next uses the **lowest** `revalidate` among a route's fetches to set the
revalidation frequency of the whole route — the segment-level
`export const revalidate` does not override a lower value inside it. Measured
against this codebase:
| Page family | Segment | Lowest fetch | Effective |
|---|---|---|---|
| `/school/[slug]` | 604800 | `fetchSchoolDetails` at 300 | **5 minutes** |
| `/schools/*` | 604800 | `fetchNationalAverages` at 3600 | **1 hour** |
The Unleash SDK polls every 15 seconds, so a flip reaches school pages within
about five minutes and place pages within the hour, unaided. Flags flip
monthly, by hand, deliberately. That is fast enough.
What this removes: two webhook integrations, a `/api/revalidate-flags` route, a
shared-secret-in-a-query-string scheme, an idempotency requirement against
duplicate and out-of-order delivery, and a rule that every fetch in
`nextjs-app/lib/` carry a cache tag. None of it has to be built, maintained, or
kept correct as new fetches are added.
**If instant flips are ever wanted**, the webhook is the way to add them, and it
is purely additive — nothing in this design has to change first.
### Two constraints this leaves behind
**Never flag content on a `force-static` page.** `app/admissions/page.tsx`
declares `export const dynamic = 'force-static'`, so it is baked at build time
and never revalidates. A flag gating anything on such a page would not take
effect until the next deploy, silently. If a flag ever needs to reach one, that
page must first move to ISR.
**A route-family flag still needs the sitemap rebuilt.** The sitemap is held in
memory and rebuilt only at startup or via `POST /api/admin/regenerate-sitemap`.
No flag in scope touches the sitemap (§6), so this is deferred with the route
case rather than solved now — but a route flag must not ship without it, or the
sitemap will advertise URLs that `notFound()`.
## 5. What "off" means, per surface
| Surface | Off |
|---|---|
| Route | `notFound()`, **and** absent from the sitemap, **and** absent from nav |
| UI element | Not rendered; surrounding page byte-identical to today |
| API field | Key **absent**, not `null` |
| API endpoint | 404, not 403 |
The three parts of the route rule move together or not at all. Submitting URLs
to Google that return 404 is the bug fixed in PR #124, and a flag is a new way
to reintroduce it.
An API field is withheld **at the source**, never rendered-but-hidden. The
precedent is already set in this codebase by commit `c9a1892`: `/api/schools/`
is public and unauthenticated, so leaving a withheld field in the payload hands
the record to anyone who opens the network tab.
## 6. First consumer: `admission_distance`
The last-distance-offered feature is merged to `main` and live on staging.
Production has never received it: `/api/schools/100010` on production carries
no `admission_distance` key, and no Distance section renders.
It needs **exactly one gate** — `backend/app.py:809`, where the field is
attached to the school payload:
```python
"admission_distance": (
supplementary.get("admission_distance")
if flags.is_enabled("admission_distance") else None
),
```
The frontend follows with no change. `DistanceSection` already returns `null`
when `admission_distance?.distance_m == null`, and `PrimarySchoolSections`
already conditions the admissions block on `(admissions || admissionDistance)`.
The off-state is the commonest state on the site — only 57 local authorities
publish cut-off distances at all — so it is well covered by construction.
The flag does not touch the sitemap: school pages exist either way.
Intended lifecycle: default off, so production receives the code dark on the
next promotion; on in the `development` environment so staging keeps testing
it; flipped on in `production` when the owner chooses.
**This flag exercises two of the three surfaces** in §5 — API field and UI
element. No route case ships with it. The route rule is specified but unproven
until a route-shaped flag exists, and should be treated as such.
## 7. Testing
**Backend unit.** The registry is well-formed; an unknown flag evaluates
`False`; `/api/flags` returns every declared flag with its default when the
SDK is unreachable; `admission_distance` is absent from the school payload when
the flag is off and present when on.
**Frontend unit.** `getFlags()` returns declared defaults when `/api/flags`
fails, rather than throwing and taking the page with it.
**E2E.** Journeys read `/api/flags` and gate flag-dependent assertions on it,
matching the `test.skip` shape the suite already uses.
One trap to avoid, worth stating because the existing distance journeys walk
straight into it: they already skip when no school has a published figure, so
with the flag off they would skip silently and the suite would go green. The
gate must be explicit — **if `/api/flags` reports `admission_distance` on, then
a school with a cut-off must be found**, converting a silent skip into a real
assertion.
## 8. Lifecycle
A flag is temporary scaffolding, and the failure mode of every flag system is
accumulation.
The registry records the date each flag was added, and a backend test fails any
flag older than **90 days**. Removing a flag means deleting the registry entry,
the branches that read it, and the flag in the Unleash UI.
Unleash SDK usage metrics stay enabled, so the UI shows which flags are still
being evaluated — the evidence needed to retire one safely.
## 9. Risks
**Production gains a homelab dependency.** If Unleash is unreachable when a
production container cold-starts with an empty cache, every flag evaluates
`False` and any feature currently switched on disappears. The fcache volume
covers restarts; the 90-day lifecycle rule bounds how long any feature is
exposed to this. It is a real regression risk and the reason flags must be
retired rather than left on indefinitely.
**Flag state is not in git.** `main` no longer tells you what production is
showing. The registry lists what *can* be flagged; only the Unleash UI says
what *is*. This is inherent to the choice of a service.
**A large promotion backlog exists.** Production is running the
pre-SEO-programme build — no place pages, and a sitemap still declaring the
apex host. The first promotion after this work ships that entire backlog. The
flag isolates the distance feature from it and nothing else.
## Out of scope
- Percentage rollouts, user targeting, A/B testing, and Unleash strategies
beyond simple on/off. Flags are booleans.
- Pipeline and dbt flags. Airflow and dbt are not flag consumers.
- Client-side flag evaluation. Flags are server-side only.
- Automatic flag removal. The staleness test reports; a person deletes.
@@ -0,0 +1,294 @@
# School Autosuggest — Design
**Date:** 2026-08-26
**Status:** approved for planning
**Depends on:** the feature-flag layer (PR #125, merged)
## Goal
Suggest schools by name as someone types in the site's main search box, so a
parent who knows the school they want reaches it in one step instead of
searching, scanning a result list, and clicking.
Scope is **schools only**. Places and postcodes were considered and excluded —
see *Out of scope*.
## The finding that shapes everything
The site's rate limiter does not do what it looks like it does.
`limiter = Limiter(key_func=get_remote_address)` with `60/minute` reads
`request.client.host`. In staging and production the backend has no published
ports and sits on the internal `backend` network, so its only caller is the
Next proxy — and `request.client.host` is therefore **the Next container**, for
every browser user on the site.
Measured against staging: 70 concurrent requests to `/api/schools` returned
**60 × 200 and 10 × 429**. One machine consumed the whole site's budget for
that minute.
Autosuggest is the worst possible feature to build on that. One person typing
"st marys primary" produces six to eight debounced requests; **eight concurrent
searchers would 429 the site.** The compare modal's search-as-you-type already
shares this bucket, so the exposure exists today — autosuggest makes it
certain.
Fixing the keying is therefore part of this work, not a follow-up.
## 1. Rate-limit keying
Both environments sit behind Cloudflare (`server: cloudflare`, `cf-ray` present
on staging and production). Cloudflare sets `CF-Connecting-IP` on every request
to the origin and **overwrites any client-supplied value**, which makes it
trustworthy in a way a parsed `X-Forwarded-For` chain is not.
```python
def client_key(request: Request) -> str:
"""Rate-limit bucket: the real caller, not the proxy in front of them."""
cf = request.headers.get("cf-connecting-ip")
if cf:
return cf.strip()
xff = request.headers.get("x-forwarded-for")
if xff:
return xff.split(",")[0].strip()
return get_remote_address(request)
```
`nextjs-app/app/api/[...path]/route.ts` already forwards every inbound header
except `host` and `connection`, so `CF-Connecting-IP` reaches the backend with
no proxy change.
**This header is trustworthy only for traffic that actually passed through
Cloudflare, and nothing in the application can verify that it did.** An earlier
draft of this section claimed Cloudflare "replaces the header, so a browser
cannot forge it", and that only the `X-Forwarded-For` fallback was forgeable.
That was wrong. Cloudflare does overwrite the header *on requests it handles* —
but a caller reaching the origin directly sets whatever it likes, and this
process cannot distinguish an edge-set header from an attacker-set one. Both
headers are equally forgeable in that scenario.
The consequence is sharper than a weakened defence. An attacker rotating
`CF-Connecting-IP` per request mints a fresh rate-limit bucket every time and
evades per-client limits entirely — including on the DataFrame-heavy
`/api/schools`. Against abuse that is *worse* than the shared bucket it
replaced, which at least capped everyone at 60/minute together.
Two mitigations, and they are not interchangeable:
1. **The real fix is at Cloudflare** — Authenticated Origin Pulls, or an origin
firewall that refuses connections not from Cloudflare's ranges. Only the
edge can vouch for its own header. This is infrastructure work and is not
part of this change; it is the thing that makes the header mean anything.
2. **The ceiling in §1.1 bounds what evading the keying can achieve** while
that remains open. It does not make the header trustworthy — it makes
trusting it survivable.
The backend being unreachable from outside the Docker network is a real second
layer, but it depends on the ingress path in front of the frontend, which this
design does not control and should not assume.
### 1.1 The ceiling, which is back
The shared bucket was acting as an accidental global throttle on a
single-process uvicorn backend that filters a 25,000-row DataFrame in-process.
Correct per-user keying removes it: the origin becomes reachable at 60/min *per
user* rather than 60/min in total, and — per above — at an unbounded rate by
anyone willing to rotate a header.
An earlier draft dropped the in-app ceiling, arguing it belonged at Cloudflare.
That argument assumed the keying was sound. It is not, so the ceiling is
load-bearing rather than redundant, and it ships here:
`GlobalRateLimitMiddleware` counts all `/api/` requests in a fixed 60-second
window against `global_rate_limit_per_minute` (3000), independent of any client
identity, and refuses with a 429 that names capacity rather than the client —
an operator has to be able to tell "one noisy client" from "the origin is
saturated". It is registered last so it is outermost: a ceiling that applies
after the expensive work has run is not a ceiling.
slowapi cannot express this. `default_limits` and `application_limits` are both
evaluated with the same `key_func`, making them per-client rather than global,
and `application_limits` only apply with `SlowAPIMiddleware` installed, which
this app does not use. Hence the explicit middleware — about thirty lines, and
obviously correct, which is what a backstop needs to be.
Requests from `127.0.0.1` are exempt. The container healthcheck runs
`curl http://localhost:80/api/data-info` from inside the container, and
starving it would fail the check, restart the container, and turn a load spike
into an outage loop. The exemption keys on the peer address, never the `Host`
header, which the caller sets.
3000/minute is an estimate, not a measurement, and worth revisiting against
real traffic.
### Per-user limits
Per-user fairness and origin protection are different jobs, and this design now
does both separately: the ceiling above for the origin, and per-route limits
for fairness. Conflating them is what produced the original behaviour, where
one bucket served the whole internet.
The existing 60/minute default is unchanged, and `/api/suggest` gets
120/minute. Both are estimates rather than measurements, and are a starting
point to revisit once the keying is correct enough for real per-user traffic to
be visible — which it was not before, because everyone shared one bucket.
## 2. `GET /api/suggest`
A dedicated endpoint, not a mode of `/api/schools`.
The existing search path calls Typesense for URNs and then filters, ranks and
sorts the full in-memory DataFrame — a pandas pass per keystroke, holding the
GIL and blocking other requests in the same worker. Suggestions need none of
it: `urn`, `school_name`, `phase`, `school_type`, `local_authority`,
`postcode` and `ofsted_rating` are all already in the Typesense document
(`pipeline/scripts/sync_typesense.py`).
```
GET /api/suggest?q=<query>&limit=8
→ 200 {"suggestions": [
{"urn": 100010, "school_name": "Brecknock Primary School",
"local_authority": "Camden", "postcode": "NW1 1AA",
"phase": "Primary", "school_type": "Community school"}
]}
```
- **Under two characters** returns `{"suggestions": []}` with 200. The
keystroke path never returns an error for ordinary input.
- **Typesense unavailable** returns `{"suggestions": []}` with 200. There is
deliberately **no DataFrame fallback**: the substring scan `/api/schools`
falls back to is precisely the cost this endpoint exists to avoid, and a
silent 25,000-row scan per keystroke is worse than no suggestions.
- **`limit` is clamped** to 20. It is a public endpoint.
- **Rate limit `120/minute`** per client, not the default 60. A 200 ms
debounce tops out near 5 requests/second while someone is actively typing,
but averages far below that across a real search; 120 leaves headroom for
bursts while still bounding one client.
- **Local authority is part of the payload, not decoration.** There are many
schools called "St Mary's"; a suggestion list without the authority is
unusable for exactly the queries autosuggest is meant to serve.
### Caching
`CACHE_RULES` gains `("/api/suggest", (60, 3600, 86400))`. Prefix queries
repeat enormously across users and school names change once a year.
The client fetch must **not** use `cache: "no-store"`. The compare modal does,
and copying that pattern would throw away both the browser cache and the ETag
304s the existing `CacheAndETagMiddleware` already provides.
Both environments currently report `cf-cache-status: DYNAMIC` — Cloudflare
ignores the `Cache-Control` the API already sends, because it does not cache
dynamic paths by default. **A Cloudflare Cache Rule for `/api/suggest*` would
let the edge absorb most of this traffic and never reach the origin.** That is
a dashboard change, it is optional, and nothing here depends on it.
## 3. The combobox
This is an ARIA combobox, not a text input with a list underneath.
**Files.** `FilterBar.tsx` is already long. The work splits three ways:
`hooks/useSchoolSuggest.ts` owns fetching, debouncing and cancellation;
`components/SuggestList.tsx` owns rendering and ARIA; `FilterBar.tsx` wires
them to the existing input and form.
**Fetching.** 200 ms debounce; minimum two characters; an `AbortController`
cancels the superseded request on every keystroke. Cancellation is not an
optimisation — without it, a slow response for `"st"` can land after the fast
one for `"st marys"` and replace a correct list with a stale one.
**Suppressed during postcode entry.** The box takes a school name *or* a
postcode, and `isValidPostcode` already distinguishes them. Suggestions do not
appear once the value parses as a postcode.
**Keyboard.** `ArrowDown`/`ArrowUp` move the active option, `Escape` closes and
keeps the typed text, `Tab` closes. `Enter` **with an option active** navigates
to that school's page. `Enter` **with none active** submits the free-text
search exactly as it does today — the existing behaviour is preserved, not
replaced.
**ARIA.** `role="combobox"` with `aria-expanded` and `aria-controls` on the
input, `aria-activedescendant` pointing at the active option, `role="listbox"`
on the list and `role="option"` on each row.
**Both instances get it.** `HomeView` renders `FilterBar` twice — hero and
sticky — from one component, so there is one implementation.
## 4. Behind a flag
Flag `school_autosuggest`, declared in `backend/flags.py`, default off.
This is the most-used control on the site and the first change to it in a
while. `app/page.tsx` is an async server component, so it reads the flag and
threads it to `FilterBar` through `HomeView` — two prop hops, explicit, no
client-side flag read.
Off means the input behaves exactly as it does today: no listener, no fetch, no
markup. Not a rendered-then-hidden dropdown.
The rate-limit keying is **not** flagged. It is a correctness fix that should
apply whether or not autosuggest is on, and flagging it would mean shipping a
known-wrong limiter into production deliberately.
## 5. Analytics
`search_submitted` already carries `via: 'input'`. Selecting a suggestion fires
it with `via: 'suggestion'` plus the chosen `urn`, so the obvious question —
does this actually help, or do people ignore it — has an answer in the data
rather than an opinion.
## 6. Testing
**Backend.** `client_key` prefers `CF-Connecting-IP`, falls back through
`X-Forwarded-For` to the remote address, and two different values get two
different buckets. `/api/suggest` returns matches, returns empty below two
characters, returns empty and 200 when Typesense is unavailable, and clamps
`limit`. That it never touches the DataFrame is asserted by making
`load_school_data` raise and requiring the endpoint to answer anyway.
**Frontend.** The hook debounces, aborts superseded requests, and drops a
late-arriving response for a stale query. The list renders the ARIA
attributes. Keyboard navigation moves the active option; `Enter` on an option
navigates; `Enter` on none submits the search.
**E2E.** With the flag on, typing a known school name shows it and selecting it
lands on that school's page. With the flag off, no combobox markup exists.
Gated on the flag the same way the distance journeys are — read the observable
effect, since `/api/flags` is denied to the public.
## 7. Risks
**Removing the accidental throttle.** Covered in §1. Correct per-user keying
means the origin is reachable at 60/minute *per user* where it was 60/minute
in total, and no in-app global cap replaces it — that job goes to Cloudflare,
which is not done as part of this change. Until it is, a determined caller
with many source addresses can put more load on a single-process origin than
they can today. Against this site's traffic that is a theoretical risk rather
than a live one, but it is a real one and it is the price of the fix.
**Cloudflare bypass — the open one.** If the origin is reachable without
passing through Cloudflare, `CF-Connecting-IP` is attacker-controlled, and
rotating it per request defeats per-client limits on every endpoint. The
ceiling in §1.1 bounds the damage to the origin's total capacity; it does not
restore per-client fairness under attack, and it cannot. Closing this properly
means Authenticated Origin Pulls or an origin firewall restricted to
Cloudflare's published ranges — infrastructure work, outside this change, and
the single most valuable follow-up here.
**Typesense becomes user-visible.** Today a Typesense outage degrades search to
a slow substring match. With autosuggest it also means the dropdown silently
stops appearing. That is the correct failure — quiet, not broken — but it makes
Typesense health worth monitoring in a way it was not before.
## Out of scope
- **Place suggestions.** The 2,646 town, authority and outcode pages are a
strong candidate and would route people onto the pages W2 built, but they
live in the place registry rather than Typesense, so it is a second index and
a ranking rule for comparing two kinds of result. Worth its own change.
- **Postcode completion.** Would put postcodes.io in the keystroke path, with
its own latency and rate limits.
- **The compare modal.** It already has search-as-you-type. Converting it to
this component is a reasonable follow-up, not part of this.
- **Recent or popular searches.** No storage for either, and no evidence yet
that they are wanted.
@@ -0,0 +1,399 @@
# Destination Measures — Design
**Date:** 2026-08-28
**Status:** awaiting review
**Scope:** secondary school detail pages only
## Goal
Say what happened to a school's leavers after they left. Two sections on the
secondary template:
- **After Year 11** — every secondary, from the KS4 destination measures
- **After the sixth form** — sixth-form schools only, from the 16-18 measures
This replaces the "Post-16 destination data coming soon" placeholder standing in
`nextjs-app/components/school/SecondaryAdmissionsSection.tsx:117` since the exam
phase taxonomy work, and fills the `ks5_destinations_pct` slot specified but
never built in `2026-07-07-exam-phase-taxonomy-design.md:201`.
Mockup, with all three data states live:
<https://claude.ai/code/artifact/5be149d6-252f-473c-9a4f-4c36b05161b0>
## The finding that shapes everything
**Suppression is per cell, and the cells sum to the cohort.**
DfE withholds a figure it considers disclosive by writing `c`. It does this at
the level of an individual destination category, not the whole school, and it
publishes the cohort total alongside. The categories form a clean partition. So
where exactly one category is suppressed, subtracting the published ones from the
cohort recovers it exactly.
Verified against three real schools in the 2022/23 file:
| School | URN | Withheld | Recovers to |
|---|---|---|---|
| North East Futures UTC | 145900 | School sixth form | **3 pupils** |
| Whitley Bay High School | 108638 | Further education | **18 pupils** |
| St Matthew's RC High School | 148389 | School sixth form | **4 pupils** |
Those are the precise numbers the `c` exists to hide, and in a random 400-school
sample **22% of mainstream secondaries** have exactly one suppressed category in
their disadvantaged group. This is the normal case, not an edge case.
Three rules follow, and everything else in this document is downstream of them.
**R1 — Never *publish* enough to derive a remainder.**
An earlier draft of this rule said "never *render* a derived remainder", and
that was the defect code review caught in PR #137. Not drawing a number does
nothing to stop it being computed: `GET /api/schools/{urn}` is public and
unauthenticated, so anything in the payload is published whatever the UI
chooses to draw. The rendering guards shipped; the payload still carried the
cohort and every published category, and `cohort - sum(published)` returned
Whitley Bay's withheld figure exactly.
The rule is therefore about the serialiser, and the UI guards are a second line
of defence behind it. Two identities have to be closed:
- within a pupil group the categories sum to the cohort, so a group with
exactly **one** suppressed category gives it away;
- across groups, disadvantaged + other = all for every category, so a category
suppressed in exactly **one** of the three gives itself away.
`_mask_for_disclosure` applies DfE's own answer — secondary suppression —
withholding a companion cell until every row and every column hides either none
or at least two. It iterates, because each new suppression can break the other
identity, and terminates because cells are only ever added.
The companion must carry pupils. Suppressing a zero looks like secondary
suppression and protects nothing: the residual still equals the original
withheld figure.
Where no companion can do the job — a sparse cohort whose every other category
is `not_applicable`, routine in special schools and alternative provision — the
pupil group is **dropped from the payload entirely**. A first version simply
returned at that point with the violation intact and no signal, which review
caught: a disclosure-control pass that fails silently is worse than none,
because everything downstream trusts it. The function now cannot terminate
except in a state where `disclosure_invariant_holds()` is true, and an
exhaustive test sweeps all 81 suppression patterns of a four-category group to
prove it.
Measured cost on the 400-school sample: the all-pupils bar survives on **94%**
of mainstream secondaries rather than 100%. That is the price of not
republishing what DfE withheld.
**R2 — Never aggregate across a suppression boundary.** Summing published
components to fill a gap is R1 with extra steps.
DfE's own aggregates (`Sustained education destination`, `Sustained education,
employment & apprenticeships`) are ingested but **not served**. An aggregate
spanning exactly one suppressed component names it, and nothing renders them
today — an unused field that leaks is not a trade-off worth carrying. They can
be re-added with their own guard if the fallback ladder is ever built.
**R3 — The three pupil groups are one disclosure surface, not three.**
Disadvantaged and Not-known-to-be-disadvantaged partition All pupils, so
rendering any *two* of them recovers the third. Where a category is suppressed in
the disadvantaged group, it must therefore also be withheld from **all other
pupils** — the all-pupils view is the primary one and keeps it.
This costs almost nothing, because DfE already applies the same masking: across
the sample, 493 of 498 suppressed disadvantaged cells were suppressed in the
other group too. The mart enforces the remaining 5, which fell on 2 schools of
262. **The all-pupils bar is unaffected** — masking the whole page wherever the
disadvantaged group is thin would remove the bar from 80% of schools, and is not
what this rule says.
R1 and R2 both hold within a group and still leak across the switch, which is why
R3 is stated separately.
### The convention that would break this quietly
`macros/safe_numeric.sql` coerces every EES sentinel — `z`, `c`, `x`, `q`, `u` —
to `NULL`, deliberately and correctly for attainment, where "suppressed" and "no
data" are equally unrenderable. Here they are not the same thing: one must print
*withheld*, the other must print nothing at all, and the difference is what keeps
R1 enforceable.
**`safe_numeric` must not be used on destination counts.** The staging model
keeps the sentinel in a companion status column. This is the single most likely
way for this feature to regress into a disclosure, so it gets its own dbt test.
## What is actually available
Measured against the EES public API (open, no key). Both datasets carry
`geographicLevel: School` with `urn` on every location option, so the join to
`dim_school` is direct.
| | KS4 | 16-18 |
|---|---|---|
| Dataset id | `019d4f41-22d1-71b2-a1a7-f3b91026815b` | `019d4e73-6440-7523-b60c-bfab1ad4a30d` |
| Rows | 1,871,739 | 3,862,658 |
| Institutions | 4,946 | 3,065 |
| Time periods | 2009/10–2022/23 | 2016/17–2022/23 |
**Destination categories (KS4).** School sixth form · Sixth form college ·
Further education · Other education destination · Sustained apprenticeships (with
level breakdown) · Sustained employment destination · Not recorded as a sustained
destination · Activity not captured. Plus the aggregates `Sustained education
destination` and `Sustained education, employment & apprenticeships`.
**16-18 adds** UK higher education institution and FE split by level, which is
what makes the post-16 section worth having.
**Breakdowns.** `Disadvantage Status` gives Disadvantaged / Not known to be
disadvantaged / Total — exactly the three-way switch. Sex, ethnicity, FSM status,
prior attainment and SEN provision also travel in the same table; we ingest none
of them.
**Indicators.** Both counts and percentages, plus the cohort size. Bar widths use
the counts — the published percentages do not sum to 100.
### Coverage, and what degrades
Random 400-school sample, 2022/23, mainstream secondaries (n=262):
| View | As published by DfE | After R1–R3 masking | Consequence |
|---|---|---|---|
| All pupils, all categories | 100% | **94%** | Bar works nearly everywhere |
| Disadvantaged, headline rate | 95% | 95% | Gap panel works |
| Disadvantaged, three grouped cards | 68% | 68% | Degrades card by card |
| Disadvantaged, all six categories | 20% | **20%** | Bar unusable for this group |
The middle column is what the site actually serves. Masking costs the
all-pupils bar on 6% of mainstream secondaries — those are schools where a
category was suppressed in exactly one pupil group and no non-zero companion
existed below the all-pupils row.
Special schools and alternative provision are far worse: 13% and 41% respectively
have the whole cohort suppressed even for all pupils. The empty state is
load-bearing, not defensive.
## The display
Question-led. Three cards over one bar, with the cards acting as a lens on the
bar rather than a summary beside it — hovering a card dims the bar, table and
England reference to the categories that card is built from. The full mockup is
linked above; what matters for implementation:
**The headline is not the sustained rate.** That figure sits between 92% and 97%
for nearly every school in England. The mix is what varies, so the mix leads.
**The grouping is ours, not DfE's.** "Academic route" = school sixth form +
sixth-form college; "College" = FE and other colleges; "Work" = apprenticeship +
employment. This is the most arguable thing on the page, so it lives in one place
in `lib/destinations.ts`, is explained in a tooltip, and is reversible in one
edit.
**The absence is hatched neutral, never a colour.** "Activity not captured" means
no record in the sources DfE holds — it includes independent schools, moving
abroad and private training. Colouring it as a bad outcome would be a factual
error rendered in CSS. The hatch also fixes a real contrast problem: neutral
against the employment blue failed CVD separation at ΔE 7.6, and texture is the
secondary encoding that rescues it. Every other adjacent pair clears ΔE 10.9
under protanopia.
**Colour tokens.** Education is one hue in three steps (school-like to
college-like); apprenticeship and employment are separate hues. Six new tokens in
`globals.css`, defined in both themes, per the existing token discipline.
**The disadvantage split rides the same control.** One visualisation serving
three cohorts, with the England reference repointing to the matching national
group. The gap statement stays visible below the bar whatever is selected,
because a gap nobody clicks on is a gap nobody sees.
## Data model
### Extraction
A new `tap-uk-ees-destinations` extractor, separate from `tap-uk-ees`. The
existing tap downloads a release ZIP and reads a CSV inside it; the destinations
files are far larger than we need and the query API filters server-side, so this
one POSTs to `/v1/data-sets/{id}/query` and pages through results.
With every dimension pinned — destination measures, disadvantage status, sex
Total, characteristic topic Total — one year returns **252,610 rows** across all
geographic levels. Three school-level years is comfortably tractable.
Pinning is mandatory, not an optimisation: leaving the characteristic dimensions
unconstrained returned 45 rows where 9 were wanted, because every breakdown
shares one table.
The tap emits the raw value as text. **It does not coerce `c`.**
### Staging
`stg_ees_ks4_destinations` / `stg_ees_ks5_destinations`. Each raw value becomes
two columns:
```sql
case when raw ~ '^-?[0-9]+(\.[0-9]+)?$' then raw::numeric end as pupils,
case
when raw ~ '^-?[0-9]+(\.[0-9]+)?$' then 'published'
when lower(trim(raw)) = 'c' then 'suppressed'
else 'not_applicable'
end as status
```
### Marts
`fact_ks4_destinations` and `fact_ks5_destinations`, **long format**:
```
urn, year, pupil_group, destination_category, cohort_pupils, pupils, percentage, status
```
This departs from the wide house pattern (`fact_ks4_performance` and friends) on
purpose. `pupil_group` is a genuine third dimension; going wide would need three
sets of every column, and R2 is far easier to test on rows than on columns.
Roughly 8 categories × 3 groups × 4,946 schools × 3 years ≈ 356k rows.
`fact_destination_national` carries the same grain for England, so the page's
England reference repoints with the switch.
### dbt tests
- `assert_destinations_no_derived_remainder` — for every (urn, year,
pupil_group) with exactly one suppressed category, assert no aggregate row
exists that would let the residual be recovered. **This is the R1 guard.**
- `assert_destinations_group_masking` — for every (urn, year, category), if the
disadvantaged group carries `suppressed`, so does the other-pupils group.
**This is the R3 guard**, applied in the mart so no consumer can reach an
unmasked combination.
- `assert_destination_status_null_agreement` — `pupils is null` wherever
`status != 'published'`, and never null where it is.
- `assert_destinations_join_dim_school` — no orphaned URNs, matching the
existing `assert_no_orphaned_facts` pattern.
## API
`GET /api/schools/{urn}` gains a `destinations` block:
```json
{
"ks4": {
"cohort_year": "2022/23",
"published": "2026-04",
"groups": {
"all": { "cohort": 180, "categories": [ … ], "aggregates": { … } },
"disadvantaged": { … },
"other": { … }
}
},
"ks5": { … }
}
```
Each category carries `pupils`, `percentage` and `status`. **The serialiser never
emits a computed remainder**, and a backend test asserts that a group containing a
suppressed category serialises no total that closes the gap.
`null` for the whole block where nothing is published — the frontend renders the
empty state from its absence, not from a sentinel.
## Frontend
| File | Kind | Job |
|---|---|---|
| `lib/destinations.ts` | pure | Category list, the academic/college/work grouping, `canAggregate()` enforcing R2, percentage derivation from counts |
| `components/school/DestinationsSection.tsx` | server | Section shell, renders **all pupils** into the HTML |
| `components/school/DestinationsView.tsx` | client | Cohort switch, card↔bar linkage |
| `components/school/Post16DestinationsSection.tsx` | server | Year 13 section, sixth-form schools only |
| `app/globals.css` | tokens | Six destination colours, both themes |
Server-first matches the directory's existing discipline — every component in
`components/school/` is a server component except `AdmissionsViewToggle`, which
is the precedent this follows. All-pupils figures are in the HTML for crawlers
and for no-JS; only the switch and the hover linkage need the client.
`lib/schoolSections.ts` gains `hasKs4Destinations` / `hasKs5Destinations` flags
and the nav items, following the existing `computeSchoolFlags` pattern.
**Placement** on the secondary template: GCSE results → After Year 11 → After the
sixth form → admissions. Destinations follow attainment because they answer "and
then what happened".
**Dating.** The latest destination year is 2022/23, published April 2026, while
the site's newest KS4 year is 2024/25. The section header states its own cohort
year, or it reads as stale data next to the GCSE section above it.
## Edge states
| State | Frequency | Behaviour |
|---|---|---|
| Whole cohort suppressed | 13% of special, 41% of AP | Section renders the explanation, no chart |
| Some categories withheld | 80% of disadvantaged views | Cards degrade individually; **no bar**; table marks withheld rows |
| Disadvantaged group suppressed entirely | 5% | Switch drops to two options, gap panel not rendered |
| No sixth form | — | Post-16 section not rendered at all — absence is correct, a "no data" placeholder would imply something is missing |
| School too new | — | "First figures expected in 2026", not a bare no |
## Testing
Per CLAUDE.md, user-facing behaviour extends `e2e/` in the same PR.
**Unit** — `lib/destinations.ts` is where R1 and R2 live, so it carries the
heaviest tests: `canAggregate()` refuses a group containing one suppressed cell,
allows one spanning two, and the bar builder refuses to emit segments for any
group with suppression. These are the tests that must fail loudly if someone
later "fixes" a gap in the chart.
**dbt** — the three tests above.
**Backend** — the serialiser emits no closing total for a partially suppressed
group.
**E2E** — a school with full data renders three cards and a bar; a school with a
partially suppressed disadvantaged group renders the withheld state and **no bar
element**; a suppressed school renders the explanation; a school with no sixth
form renders no post-16 section.
Note the staging caveat: mart changes are inert until the Airflow pipeline runs,
and the staging E2E gate runs post-merge.
## Out of scope
- **Compare view and rankings.** The long mart shape supports both; neither is
built here. Flagged because "% to a school sixth form" is a plausible rankings
metric and the mart shape should not have to change to allow it.
- **Ethnicity, sex, SEN and prior-attainment breakdowns.** Available in the same
file, ingested deliberately not at all — each is a separate editorial decision
about what a school page should assert.
- **Longer term destinations** (3 and 5 years out) and **Progression to higher
education** — separate publications, worth a later look for sixth forms.
- **Primary schools.** No KS2 destination measures publication exists; DfE
tracking starts at KS4. Naming the secondaries a primary's leavers go to needs
the National Pupil Database, which is not publishable at that grain.
## Risks
**A later change reintroduces the disclosure.** The likeliest routes are
applying `safe_numeric` to a destination column for consistency, adding a
`coalesce` in a mart, or — as happened in review — enforcing a disclosure rule
at the rendering layer instead of the publishing layer. Mitigation is the dbt
tests plus `backend/tests/test_destinations_api.py`, which reconstructs the
residual the way an attacker would and asserts it no longer resolves.
**The two-year lag reads as staleness.** Mitigated by dating the cohort in the
section header rather than only in a tooltip.
**Sixth-form retention will be misread.** "41% went to a school sixth form" says
nothing about *which* school. The published file reports destination type, never
destination institution. Copy must never imply "stayed on here", and the tooltip
should say so.
**Section length.** The secondary template is already long and this adds two
sections. If it becomes a problem the post-16 section is the one to collapse
behind a disclosure, not the Year 11 one.
## Open questions
1. Is the disadvantage split its own section or a sub-block inside the
destinations section? Modelled as a sub-block; it is the most differentiating
figure on the page and the most easily misread on a small cohort.
2. Do we ingest the apprenticeship level breakdown (intermediate / advanced /
higher) now, or collapse to one apprenticeship figure and revisit? Collapsed
in this design.
@@ -0,0 +1,368 @@
# Giving schoolcompare a human author: an About page and a blog
**Date:** 2026-09-02
**Status:** Design — awaiting review
**Scope:** A named author for the site, an `/about` page, and a Payload-CMS-backed
blog at `/blog`.
## Why
The site reads as synthetic. Not because of its tone, but because of three
specific absences:
1. **Nobody is accountable for the numbers.** There is no author, no statement
of why the site exists, and no one who can be wrong. The only human trace on
the entire site is `contact@schoolcompare.co.uk` in the footer.
2. **No visible judgement.** Every figure is presented as though it fell out of
a machine. Hundreds of editorial decisions went into this codebase — which
metrics to show, when a benchmark is invalid, what to suppress — and not one
of them is visible to a reader. `isSpecialSchool()` silently drops the
England comparison for special schools and PRUs because that comparison is
meaningless; nowhere does the site *say* so.
3. **The voice is institutional third person.** "schoolcompare brings it all
into one place." "Built for parents, governors, journalists." That is
brochure register, and it is precisely the register that machine-generated
content defaults to.
There is a second, independent reason. The SEO programme
(`2026-08-20-seo-programme-design.md`) defines eight workstreams and none of
them address E-E-A-T or authorship. School performance data is YMYL territory;
an anonymous site republishing DfE figures has no authorship signal at all. This
work fills that hole, and the blog gives W6 (explainer content) somewhere to
live.
### The failure mode to avoid
The standard fix — a stock photo and "Hi, I'm Tudor, and I'm passionate about
education!" — reads as *more* synthetic than the current coldness. Manufactured
warmth is a stronger machine-tell than plain institutional voice. Everything
here has to be specific, occasionally awkward, and willing to be unflattering,
or it makes the problem worse.
## Positioning
The author is **Tudor**: first name only, real photograph, no surname, no
employer named.
The credibility claim is deliberately **not** educational expertise. The About
page states plainly: *"I'm not an education expert."* Authority comes from two
things that are actually true:
- **Experience.** A parent going through primary admissions in south-west London
right now. Google's E-E-A-T leads with Experience, and lived experience of the
thing is exactly what the DfE's own service lacks.
- **Method.** Every number's provenance is stated, so a reader can check the
site rather than trust it.
This is more durable than borrowed expertise: it cannot be undermined by someone
noticing the author has no teaching qualification.
**Consequence for the design.** A `Person` entity with no surname is a weak
search signal and cannot be corroborated off-site. The credibility load
therefore shifts onto the methodology being visibly rigorous. That is a design
constraint, not a caveat — it is why the About page carries a substantial
"how this is built and where it can be wrong" section rather than a short bio.
### Voice rules
Applied to About and every post. Recorded here so the voice does not drift.
- First person singular. "I built", not "we provide".
- Concrete over general. "when we were looking at schools in Wandsworth" beats
any amount of stated warmth.
- State limits before someone else finds them. Every post that presents a
metric says what it does not show.
- No mission statements, no "passionate about", no invented team.
- No em dashes. One of the clearest tells of machine-written prose, which is
the exact problem this work exists to fix.
- Short sentences. The existing code comments in this repo are already written
this way; the prose should match.
## Scope
**In:**
- `/about` — a coded page (not CMS-managed).
- `/blog` and `/blog/[slug]` — Payload-backed, with an index and post pages.
- Payload CMS installed into the existing Next application.
- Footer and navigation links to both.
- `Person`, `Organization`, `BlogPosting`, `BreadcrumbList` JSON-LD.
- RSS feed and sitemap integration.
- One first post, so the blog does not launch empty.
**Out (deliberately):**
- Rewriting existing homepage/how-it-works copy into first person. Worth doing,
but it would double the review surface of this PR. Separate change.
- In-product signed notes on school pages (the "distributed humanity" idea).
Revisit once About and the blog exist.
- Comments, newsletter, author accounts beyond one.
- A team page. There is no team.
## Architecture
### Topology
Payload 3 installs **into the existing Next application** and serves `/admin`
from the same container. One image, one deploy, no new service. This is
Payload 3's native model and it makes on-demand revalidation trivial, because
the CMS hooks run in the same process as the Next cache.
Accepted costs: the public site's image now carries Payload, so a CMS security
patch redeploys the whole site; and the image grows substantially.
### Two collisions that must be handled
**1. `/api` is already taken.** `app/api/[...path]/route.ts` is a catch-all that
proxies `/api/*` to FastAPI at runtime. Payload's default API route is also
`/api`. Left alone, these fight, and the failure is not clean — the catch-all
would swallow Payload's admin API calls and forward them to FastAPI.
Payload's API route is therefore remapped:
```ts
routes: { api: '/cms-api', admin: '/admin' }
```
with its route group at `app/(payload)/cms-api/[...slug]/route.ts`. The
`/cms-api` prefix must also be added to the FastAPI proxy's excluded-paths list
as a defensive second line.
**2. `next.config.js` is CommonJS.** Payload's `withPayload()` wrapper is ESM
only. The config must become `next.config.mjs`, converting `module.exports` to
`export default` and wrapping the export. All existing content — the standalone
output, `outputFileTracingIncludes`, the staging `X-Robots-Tag` header block,
the CSP — carries over unchanged. This is mechanical but it touches the file
that controls staging's noindex, so it needs care and an explicit test.
### Database
Payload uses the existing `sc_database` Postgres instance, in its **own
`payload` schema**:
```ts
db: postgresAdapter({
pool: { connectionString: process.env.DATABASE_URL },
schemaName: 'payload',
})
```
The frontend container is already on the `backend` Docker network, so it can
reach `sc_database:5432` with no networking change. It needs a new
`DATABASE_URL` environment variable.
Schema isolation is not cosmetic. `public` currently holds the application
tables and Airflow's metadata, and `scripts/migrate_csv_to_db.py --drop` exists
to drop and reimport. Blog content living in its own schema means no data
pipeline operation can destroy it.
**Verified 2026-09-02** (this was an open question when the spec was written).
`--drop` calls `run_full_migration()` in `backend/migration.py`, which drops
exactly two tables by name:
```python
ks2_tables = ["school_results", "schools"]
for tname in ks2_tables:
if tname in existing:
Base.metadata.tables[tname].drop(bind=engine)
```
There is no `Base.metadata.drop_all()` anywhere in `backend/`, and no
`DROP SCHEMA`. The only other drop is `_apply_schema_drops()`, a single
schema-qualified `DROP TABLE IF EXISTS marts.fact_parent_view CASCADE`.
Nothing sets `search_path`, so the SQLAlchemy metadata resolves to `public`,
and `inspector.get_table_names()` does not even enumerate other schemas.
So the guarantee is stronger than schema isolation alone: `--drop` targets two
named tables that Payload does not have, and would not reach `posts`, `media`
or `users` even if they shared a schema. The `payload` schema remains the right
choice — it protects against a *future* broadening of that script rather than
today's behaviour — but the safety claim rests on verified code, not on
assumption.
Putting CMS tables in this instance is consistent with existing practice —
Airflow already stores its metadata there.
### Migrations
Payload's Postgres adapter auto-pushes schema in development and requires
explicit migrations in production. Use `prodMigrations`, which runs pending
migrations during server initialisation:
```ts
db: postgresAdapter({ /* ... */, prodMigrations: migrations })
```
This is preferred over a one-shot init container (the `airflow-init` pattern)
because the app is a single long-running process and there is no ordering
problem to solve. Migration files are generated with `payload migrate:create`
and committed, so schema changes travel through the same PR and staging gate as
code.
### Media
Uploads go to a Docker named volume, consistent with `postgres_data`,
`typesense_data` and `airflow_logs`.
- `staticDir` must be an **absolute** path in Payload 3: `/app/media`.
- The container runs as `nextjs` (uid 1001). The Dockerfile must
`mkdir -p /app/media && chown nextjs:nodejs /app/media` **before** the volume
is mounted, or Docker will create the mountpoint root-owned and every upload
will fail with EACCES.
- `sharp` moves from `devDependencies` to `dependencies` — Payload needs it at
runtime to generate `imageSizes`.
- The volume must be added to the backup routine alongside Postgres. A blog
post's images are not reproducible from the pipeline.
### Rendering
**Constraint:** CI builds the image with no database reachable. Blog pages
therefore cannot use build-time `generateStaticParams` — that would either fail
the build or bake in an empty post list.
Instead: ISR. Post and index pages declare a `revalidate` window and render on
first request, with Payload `afterChange` / `afterDelete` hooks calling
`revalidatePath('/blog')` and `revalidatePath('/blog/' + slug)` for immediate
publication. Because Payload runs in the same process, the hook calls
`revalidatePath` from `next/cache` directly — no webhook, no shared secret.
The ISR cache lives on container disk and is cleared by a redeploy. For a
single container serving a handful of posts this is fine.
### Collections
- **`posts`** — `title`, `slug`, `publishedAt`, `excerpt`, `heroImage`
(relation to `media`), `content` (Lexical rich text), `seo` group
(`metaTitle`, `metaDescription`), `_status` (drafts enabled).
- **`media`** — upload collection, `alt` required, `imageSizes` for thumbnail
and hero widths, public read access.
- **`users`** — Payload's auth collection. One account. Public creation
disabled.
Drafts are enabled so posts can be written over several sittings and previewed
before publication.
**Payload Blocks** are how posts embed live product components — a real trend
chart or comparison table inside a post, rendered from live data rather than
screenshotted. This is the main thing the CMS has to earn back against
file-based MDX, and it directly serves the goal: showing judgement in context.
Ship with one block (a callout/aside for "what this number doesn't tell you");
add a live-chart block once a post needs it.
### Security
`/admin` is the first authenticated surface on this site. Public, hardened:
- `PAYLOAD_SECRET` — long, random, set in the Portainer stack environment, never
committed. The same variable must exist in staging with a *different* value.
- Strong unique password on the single admin account.
- Login rate limiting via Payload's `maxLoginAttempts` / `lockTime`.
- `X-Robots-Tag: noindex, nofollow` on `/admin/*` and `/cms-api/*`, and a
`robots.ts` disallow. The admin panel must never be indexed.
- Public user creation disabled; no open registration.
- Verify the existing CSP `frame-ancestors` directive does not break the admin
panel.
Residual risk, accepted: a future Payload authentication CVE is live against the
public internet. Mitigation is prompt patching, which the staging→prod pipeline
already supports. If this becomes uncomfortable, restricting `/admin` at the
proxy to LAN/VPN is a one-line change later.
Staging note: staging runs the same image on `stx.`, so it gets its own admin
panel and its own database. It must have its own `PAYLOAD_SECRET` and its own
credentials — never production's.
## Deployment changes
- `nextjs-app/Dockerfile` — create and chown `/app/media`; ensure Payload's
admin bundle and `sharp` survive standalone output file tracing.
- `docker-compose.portainer.yml` and the staging equivalent — add
`DATABASE_URL` and `PAYLOAD_SECRET` to the `frontend` service, add a
`payload_media` volume mounted at `/app/media`, and add
`depends_on: sc_database`.
- Document both new environment variables in the compose header comment block,
which is where this stack records its configuration.
## SEO
- `Person` (Tudor, with photo) and `Organization` JSON-LD on `/about`.
- `BlogPosting` + `BreadcrumbList` on post pages, with `author` referencing the
same `Person`.
- Canonical URLs on `/blog` and every post.
- Posts and `/about` added to the existing sitemap (`app/sitemap.xml/route.ts`
and `app/sitemaps/[...parts]`). Post URLs come from Payload at request time.
- RSS feed at `/blog/rss.xml`.
- Footer links to both pages, under a new "About" column.
**Navigation is deliberately left alone.** `Navigation.tsx` renders a bottom tab
bar on mobile that already carries four items (Search, Compare, Rankings,
Admissions). A fifth tab makes each one cramped at 320px, and About and Blog are
both lower-intent than any of the four. Both live in the footer; About
additionally gets a byline link from every post, which is where a reader who
cares actually asks the question. Revisit only if analytics show people hunting
for it.
## Testing
Unit (Jest):
- Post rendering, including a post with no hero image and one with no excerpt.
- Slug generation and collision handling.
- JSON-LD shape for `BlogPosting` and `Person`.
- The `next.config.mjs` conversion preserves the staging `X-Robots-Tag` rule —
this guards the riskiest mechanical change in the plan.
E2E (Playwright, `e2e/`, required by CLAUDE.md for user-facing change):
- `/about` renders, shows the author name and photo, and is reachable from the
footer and nav.
- `/blog` lists at least one post; clicking through reaches the post.
- A post page renders title, date, body and byline.
- `/admin` responds with `noindex` and does not leak a stack trace when
unauthenticated.
Note the known constraint: new journeys cannot be proven in PR checks, because
the staging E2E gate runs post-merge.
## Risks
| Risk | Mitigation |
|---|---|
| `next.config.mjs` conversion silently drops the staging noindex header, making staging a crawlable duplicate | Unit test asserting the header rule; verify on staging before promotion |
| Payload API route collides with the FastAPI `/api` proxy | Remap to `/cms-api`; add to the proxy's exclusion list |
| Media volume mounts root-owned; all uploads fail with EACCES | `mkdir`+`chown` in the Dockerfile before the mount; test an upload on staging |
| Build fails or bakes empty content because CI has no DB | No build-time DB access; ISR only |
| A pipeline `--drop` destroys blog content | Separate `payload` schema; verify `--drop` blast radius before building |
| Media volume not backed up; images unrecoverable | Add `payload_media` to the backup routine |
| Payload auth CVE exposed publicly | Prompt patching; proxy restriction available as a fallback |
| Blog launches empty or goes stale | Ship with one post; cadence is explicitly "a few times a year", so no cadence is promised anywhere on the page — no dates implying a schedule |
## Sequence
Each step is independently reviewable and mergeable.
1. **Payload foundation** — install, `next.config.mjs` conversion, `payload`
schema, `/cms-api` remap, `users` collection, `/admin` hardening, compose and
Dockerfile changes. No public-facing change yet. Verify on staging that the
site is unchanged and `/admin` works.
2. **`/about`** — coded page, photo, `Person`/`Organization` JSON-LD, footer and
nav links, e2e journey. Independently valuable and does not depend on the
blog.
3. **Blog** — `posts` and `media` collections, `/blog` index and post pages, ISR
plus revalidation hooks, RSS, sitemap, structured data, e2e journeys.
4. **First post** — written in the admin panel, published through the normal
flow, proving the whole path end to end.
Step 1 carries all the infrastructure risk and none of the visible benefit, so
it should be verified on staging carefully before step 2 starts.
## Dependencies on Tudor
- **A photograph.** Blocks step 2. Nothing else in the plan is blocked by it.
- **The first post's subject.** Blocks step 4 only. Suggested: what school
performance data cannot tell you — it demonstrates judgement, is genuinely
useful, and is the kind of thing an anonymous or machine-written site will not
publish.
- ~~Confirmation that `scripts/migrate_csv_to_db.py --drop` is schema-scoped.~~
**Resolved 2026-09-02** — verified in `backend/migration.py`; see the
Database section. No action needed.
+795 -3
View File
@@ -1226,6 +1226,27 @@ const CUTOFF_CANDIDATE_URNS = [
101099, 100553, 102574, 100769, // mixed 101099, 100553, 102574, 100769, // mixed
]; ];
/**
* Whether the last-distance-offered feature is switched on here.
*
* Read from the data rather than from /api/flags, which the public proxy
* denies on purpose — the endpoint names unreleased features. The observable
* effect is the field's presence: the flag is off iff no candidate school
* carries an `admission_distance` key at all.
*
* The distinction that matters: `admission_distance: null` means this school
* has no published cut-off, and the key being ABSENT means cut-offs are not
* being published at all.
*/
async function distanceFeatureIsOn(page: Page): Promise<boolean> {
for (const urn of CUTOFF_CANDIDATE_URNS) {
const res = await page.request.get(`/api/schools/${urn}`);
if (!res.ok()) continue;
if ('admission_distance' in (await res.json())) return true;
}
return false;
}
async function schoolWithCutoff(page: Page) { async function schoolWithCutoff(page: Page) {
for (const urn of CUTOFF_CANDIDATE_URNS) { for (const urn of CUTOFF_CANDIDATE_URNS) {
const res = await page.request.get(`/api/schools/${urn}`); const res = await page.request.get(`/api/schools/${urn}`);
@@ -1237,6 +1258,105 @@ async function schoolWithCutoff(page: Page) {
return null; return null;
} }
test('when the distance feature is on, a school with a cut-off is findable', async ({ page }) => {
/*
* The gate that stops the other distance journeys passing vacuously.
*
* They all skip when schoolWithCutoff() finds nothing, which is right when
* the feature is off — but it means a feature that is *supposed* to be on
* and is silently broken shows up as a green run full of skips. This test
* fails in that case.
*/
test.skip(!(await distanceFeatureIsOn(page)),
'the admission_distance flag is off in this environment');
expect(await schoolWithCutoff(page),
'the distance feature is on, but no candidate school has a cut-off — '
+ 'the flag is on and the data or the query behind it is broken')
.not.toBeNull();
});
test('with the distance feature off, the section is absent rather than empty', async ({ page }) => {
// Shipping dark means the page renders as it did before the feature existed,
// not as a feature with its content removed.
test.skip(await distanceFeatureIsOn(page),
'the admission_distance flag is on in this environment');
// A school that exists, found rather than hardcoded — a 404 page would
// satisfy the absent-heading assertion without proving anything.
//
// A plain loop, not Array.find: find's predicate is synchronous, so an async
// one returns a Promise, every Promise is truthy, and it would always hand
// back the first URN whether or not that school exists.
let urn: number | null = null;
for (const candidate of CUTOFF_CANDIDATE_URNS) {
if ((await page.request.get(`/api/schools/${candidate}`)).ok()) {
urn = candidate;
break;
}
}
expect(urn, 'no candidate school resolves in this environment').not.toBeNull();
await page.goto(`/school/${urn}`);
await expect(page.locator('h1')).toBeVisible();
await expect(page.getByRole('heading', { name: /How far away are you\?/ }))
.toHaveCount(0);
});
/**
* A secondary school carrying an EES admissions row, which is what makes its
* Admissions section render while the distance feature is dark.
*/
async function secondarySchoolWithAdmissions(page: Page) {
const list = await page.request.get('/api/schools?phase=secondary&page_size=40');
if (!list.ok()) return null;
const body = await list.json();
for (const s of (body?.schools ?? []).slice(0, 25)) {
const res = await page.request.get(`/api/schools/${s.urn}`);
if (!res.ok()) continue;
const detail = await res.json();
if (detail?.admissions == null) continue;
return { urn: s.urn as number };
}
return null;
}
test('with the distance feature off, a secondary page makes no claim about publication', async ({ page }) => {
/*
* Shipping dark must not put words in the council's mouth. The secondary
* template is the only one that words the absence, and "X has not published
* a cut-off distance for this school" is false wherever X does publish and
* we are simply withholding it.
*
* This is why the API omits the key rather than sending null: absent means
* "cut-offs are not published at all", null means "this school has none".
* Only the second is a fact about the school, and only the second is sayable.
*/
test.skip(await distanceFeatureIsOn(page),
'the admission_distance flag is on in this environment');
const found = await secondarySchoolWithAdmissions(page);
test.skip(found === null, 'no secondary school in the sample has an admissions row');
await page.goto(`/school/${found!.urn}`);
await expect(page.locator('h1').first()).toBeVisible({ timeout: 15_000 });
// The Admissions section is still there — this is not a test that the whole
// section vanished, which would pass for the wrong reason.
await expect(page.locator('#admissions')).toHaveCount(1);
await expect(page.getByText(/has not published a cut-off distance/)).toHaveCount(0);
await expect(page.getByText(/Contact the admissions authority/)).toHaveCount(0);
});
test('/api/flags is not reachable from the public internet', async ({ page }) => {
// It names every unreleased feature and whether it is on. Next reads it
// server-side over the Docker network; the public proxy must deny it.
const res = await page.request.get('/api/flags');
expect(res.status()).toBe(404);
});
test('a published cut-off distance is shown with the year it belongs to', async ({ page }) => { test('a published cut-off distance is shown with the year it belongs to', async ({ page }) => {
const found = await schoolWithCutoff(page); const found = await schoolWithCutoff(page);
test.skip(found === null, 'no school in the sample has a published cut-off distance yet'); test.skip(found === null, 'no school in the sample has a published cut-off distance yet');
@@ -1649,13 +1769,29 @@ const CANONICAL_ROUTES: Array<[string, string]> = [
['/admissions', 'https://www.schoolcompare.co.uk/admissions'], ['/admissions', 'https://www.schoolcompare.co.uk/admissions'],
]; ];
/**
* Next normalises canonical URLs against `trailingSlash: false`, so the root
* ships as `https://www.schoolcompare.co.uk` with no slash while every other
* route keeps its path. Both forms address the same document, and which one
* Next emits is its business, not something worth pinning a test to.
*
* The first cut hardcoded the slash and failed only on the homepage — the
* same gap as the doubled brand: it asserted the metadata object rather than
* what the page actually renders.
*/
function sameUrl(a: string | null, b: string): boolean {
const strip = (u: string) => u.replace(/\/+$/, '');
return strip(a ?? '') === strip(b);
}
for (const [path, expected] of CANONICAL_ROUTES) { for (const [path, expected] of CANONICAL_ROUTES) {
test(`${path} declares exactly one canonical, on the www host`, async ({ page }) => { test(`${path} declares exactly one canonical, on the www host`, async ({ page }) => {
await page.goto(path); await page.goto(path);
const hrefs = await page.locator('link[rel="canonical"]').evaluateAll( const hrefs = await page.locator('link[rel="canonical"]').evaluateAll(
(els) => els.map((e) => e.getAttribute('href'))); (els) => els.map((e) => e.getAttribute('href')));
expect(hrefs, `${path} should declare one canonical`).toHaveLength(1); expect(hrefs, `${path} should declare one canonical`).toHaveLength(1);
expect(hrefs[0]).toBe(expected); expect(sameUrl(hrefs[0], expected),
`${path} canonical was ${hrefs[0]}, expected ${expected}`).toBe(true);
}); });
} }
@@ -1663,7 +1799,8 @@ test('a filtered homepage still canonicalises to the bare root', async ({ page }
await page.goto('/?search=primary&phase=primary&sort=name&page=2'); await page.goto('/?search=primary&phase=primary&sort=name&page=2');
const href = await page.locator('link[rel="canonical"]').first() const href = await page.locator('link[rel="canonical"]').first()
.getAttribute('href'); .getAttribute('href');
expect(href).toBe('https://www.schoolcompare.co.uk/'); expect(sameUrl(href, 'https://www.schoolcompare.co.uk/'),
`filtered homepage canonical was ${href}`).toBe(true);
}); });
test('a school page canonicalises to its own slug on the www host', async ({ page }) => { test('a school page canonicalises to its own slug on the www host', async ({ page }) => {
@@ -1714,10 +1851,33 @@ test('staging answers noindex, and stays crawlable so the noindex is seen', asyn
// The other half, and the reason this is one test rather than two: a // The other half, and the reason this is one test rather than two: a
// Disallow would stop Google fetching the page at all, so it would never // Disallow would stop Google fetching the page at all, so it would never
// see the noindex above. The two only work together. // see the noindex above. The two only work together.
//
// Scoped to the `*` group. The first cut matched `Disallow: /` anywhere in
// the file and tripped over the AI-crawler groups Cloudflare injects —
// ClaudeBot, GPTBot, Amazonbot and friends all carry a blanket disallow,
// deliberately, and none of them is Googlebot.
const robots = await (await page.request.get('/robots.txt')).text(); const robots = await (await page.request.get('/robots.txt')).text();
expect(robots).not.toMatch(/^\s*Disallow:\s*\/\s*$/mi); expect(blocksEverything(robots, '*'),
'the * group must not disallow the whole site, or the noindex is never seen')
.toBe(false);
}); });
/** True when `agent`'s group in a robots.txt disallows the entire site. */
function blocksEverything(robots: string, agent: string): boolean {
let current: string | null = null;
let blocked = false;
for (const raw of robots.split('\n')) {
const line = raw.split('#')[0].trim();
if (!line) continue;
const [key, ...rest] = line.split(':');
const value = rest.join(':').trim();
const k = key.trim().toLowerCase();
if (k === 'user-agent') current = value;
else if (current === agent && k === 'disallow' && value === '/') blocked = true;
}
return blocked;
}
test('a school page on staging is noindexed too, not just the homepage', async ({ page }) => { test('a school page on staging is noindexed too, not just the homepage', async ({ page }) => {
const list = await page.request.get('/api/schools?search=primary&per_page=1'); const list = await page.request.get('/api/schools?search=primary&per_page=1');
const [first] = (await list.json()).schools ?? []; const [first] = (await list.json()).schools ?? [];
@@ -1864,11 +2024,115 @@ test('a place page links its phase variants, and they resolve', async ({ page })
await expect(page.locator('h1')).toContainText(new RegExp(`${phase} schools in`, 'i')); await expect(page.locator('h1')).toContainText(new RegExp(`${phase} schools in`, 'i'));
}); });
/*
* The table shipped with one column of scores. A parent shortlisting from a
* town page needs to know whether a school takes their child's age, whether
* it is a faith school, and — for a primary — whether it has a nursery,
* before a percentage means anything.
*
* These assert the column headings rather than the values: nursery_provision
* and parliamentary_constituency are optional mart columns, and on an
* environment whose pipeline has not rebuilt them the API degrades them to
* absent. A value assertion would then fail for a data reason, not a code one.
*/
async function phasedPlace(page: Page, phase: 'primary' | 'secondary') {
const place = await firstPlaceOfKind(page, 'town');
const detail = await (await page.request.get(`/api/places/town/${place.slug}`)).json();
test.skip(!(detail.place.phases ?? []).includes(phase),
`no ${phase} page clears the threshold here`);
return place;
}
test('a primary place page names each school as well as scoring it', async ({ page }) => {
const place = await phasedPlace(page, 'primary');
await page.goto(`/schools/${place.slug}/primary`);
for (const heading of ['Ages', 'Religious character', 'Nursery', 'Constituency']) {
await expect(page.getByRole('columnheader', { name: heading, exact: true }))
.toBeVisible();
}
// age_range rides in on SCHOOL_COLUMNS and predates the optional columns,
// so it is the one attribute safe to assert a value for anywhere.
await expect(page.locator('table tbody td').filter({ hasText: /^\d+–\d+$/ }).first())
.toBeVisible();
});
test('a secondary place page does not ask about nurseries', async ({ page }) => {
const place = await phasedPlace(page, 'secondary');
await page.goto(`/schools/${place.slug}/secondary`);
await expect(page.getByRole('columnheader', { name: 'Ages', exact: true }))
.toBeVisible();
await expect(page.getByRole('columnheader', { name: 'Nursery', exact: true }))
.toHaveCount(0);
});
test('the measure stays beside the school name, not behind a swipe', async ({ page }) => {
// Six columns overflow a phone; .tableWrap turns that into a horizontal
// scroll. With the measure last, the number the page exists for is the one
// off the screen.
const place = await phasedPlace(page, 'primary');
await page.setViewportSize({ width: 390, height: 844 });
await page.goto(`/schools/${place.slug}/primary`);
const second = page.locator('table thead th').nth(1);
await expect(second).toContainText(/reading, writing/i);
await expect(second).toBeInViewport();
});
test('phase variants are submitted in the places sitemap', async ({ page }) => { test('phase variants are submitted in the places sitemap', async ({ page }) => {
const xml = await (await page.request.get('/sitemaps/places-1.xml')).text(); const xml = await (await page.request.get('/sitemaps/places-1.xml')).text();
expect(xml).toMatch(/\/schools\/[a-z0-9-]+\/primary</); expect(xml).toMatch(/\/schools\/[a-z0-9-]+\/primary</);
}); });
test('authority phase variants are submitted, and in their own namespace', async ({ page }) => {
// 302 of these were in the sitemap for weeks and every one 404'd: the spec
// called for the route, the plan built the bare authority page and dropped
// it, and the sitemap — written from the registry — kept submitting them.
const xml = await (await page.request.get('/sitemaps/places-1.xml')).text();
expect(xml).toMatch(/\/schools\/authority\/[a-z0-9-]+\/primary</);
});
test('every place link a place page emits resolves', async ({ page }) => {
/*
* The guard that was missing. Each family built its own links, so a URL
* shape belonging to one namespace was used by all four: an authority page
* offered "Primary schools in Barnet" pointing at /schools/barnet/primary,
* the *town*. For 87 of 151 authorities that 404'd; for the other 64 it
* quietly served a different set of schools under the same name.
*
* Only /schools links are followed. The per-school links are the same
* component the school-page journeys already cover, and there are hundreds
* of them on a page.
*/
for (const kind of ['town', 'authority', 'outcode'] as const) {
const place = await firstPlaceOfKind(page, kind);
const prefix = kind === 'authority' ? '/schools/authority/'
: kind === 'outcode' ? '/schools/near/' : '/schools/';
await page.goto(`${prefix}${place.slug}`);
const hrefs = [...new Set(
await page.locator('a[href^="/schools"]').evaluateAll(
(els) => els.map((e) => e.getAttribute('href')!)))];
expect(hrefs.length, `${kind} page links no other place`).toBeGreaterThan(0);
for (const href of hrefs) {
const res = await page.request.get(href);
expect(res.status(), `${kind} page links ${href}`).toBe(200);
}
}
});
test('an outcode page offers no phase link, because no such page exists', async ({ page }) => {
// Nobody searches "primary schools in SW11", so the spec gives outcodes no
// phase route. The registry computed the variants anyway and the page
// linked them, putting two 404s on each of 1,720 outcode pages.
const place = await firstPlaceOfKind(page, 'outcode');
const detail = await (await page.request.get(
`/api/places/outcode/${place.slug}`)).json();
expect(detail.place.phases).toEqual([]);
await page.goto(`/schools/near/${place.slug}`);
await expect(page.getByRole('navigation', { name: 'By phase' })).toHaveCount(0);
});
test('no page title repeats the brand', async ({ page }) => { test('no page title repeats the brand', async ({ page }) => {
// The root layout appends '| schoolcompare' to a plain-string title. Any // The root layout appends '| schoolcompare' to a plain-string title. Any
// route whose title already carries the brand must opt out with // route whose title already carries the brand must opt out with
@@ -1885,3 +2149,531 @@ test('no page title repeats the brand', async ({ page }) => {
expect(brands, `${path} repeats the brand: ${title}`).toBeLessThanOrEqual(1); expect(brands, `${path} repeats the brand: ${title}`).toBeLessThanOrEqual(1);
} }
}); });
test('a place straddling a boundary names every authority it sits in', async ({ page }) => {
// A quarter of outcodes and a third of towns cross an authority boundary —
// SW19 is mostly Merton but partly Wandsworth. Naming only the largest
// asserts something false about the place.
const { places } = await (await page.request.get('/api/places')).json();
const outcode = places.find((p: { kind: string }) => p.kind === 'outcode');
expect(outcode).toBeTruthy();
// Find any place the registry reports as straddling.
let straddling: { kind: string; slug: string } | null = null;
for (const p of places.filter((p: { kind: string }) => p.kind === 'outcode').slice(0, 40)) {
const d = await (await page.request.get(`/api/places/outcode/${p.slug}`)).json();
if ((d.place.authorities ?? []).length > 1) { straddling = p; break; }
}
test.skip(!straddling, 'no straddling outcode found in the sample');
const detail = await (await page.request.get(
`/api/places/outcode/${straddling!.slug}`)).json();
await page.goto(`/schools/near/${straddling!.slug}`);
for (const a of detail.place.authorities) {
if (a.slug) {
await expect(page.locator(`a[href="/schools/authority/${a.slug}"]`).first())
.toBeVisible();
} else {
// No page of its own — City of London and the Isles of Scilly are
// under the threshold. Named, deliberately not linked.
await expect(page.locator('header p')).toContainText(a.name);
await expect(page.getByRole('link', { name: a.name })).toHaveCount(0);
}
}
});
test('a place page lists its schools alphabetically', async ({ page }) => {
// Someone on a place page is usually looking for a school they can name,
// so the order should serve scanning for it. /rankings is where the
// league-table ordering lives.
const { places } = await (await page.request.get('/api/places')).json();
const town = places.find((p: { kind: string; count: number }) =>
p.kind === 'town' && p.count >= 5);
expect(town).toBeTruthy();
await page.goto(`/schools/${town.slug}`);
/*
* Per table, not per page.
*
* An unphased place page renders one table per phase, and an all-through
* school legitimately appears in both — so the page's school links are not
* one alphabetical run and never were. This assertion used to collect them
* all together and only passed because no town it picked happened to hold an
* all-through school; when the data gave Abbots Langley one, Breakspeare
* School showed up in the primary table and again in the secondary, and the
* test failed on correct behaviour.
*/
const tables = page.locator('table');
const tableCount = await tables.count();
expect(tableCount).toBeGreaterThan(0);
let checked = 0;
for (let i = 0; i < tableCount; i++) {
const names = await tables.nth(i).locator('a[href^="/school/"]').allTextContents();
if (names.length < 2) continue; // a one-row table says nothing about order
const sorted = [...names].sort((a, b) =>
a.toLowerCase().localeCompare(b.toLowerCase()));
expect(names, `table ${i + 1} is not alphabetical`).toEqual(sorted);
checked++;
}
expect(checked, 'no table had enough rows to check the ordering').toBeGreaterThan(0);
});
test('the rankings page still orders by score, not name', async ({ page }) => {
// Alphabetical is a place-page decision, not a site-wide one.
const res = await page.request.get('/api/rankings?metric=rwm_expected_pct&phase=primary');
expect(res.ok()).toBeTruthy();
const scores = ((await res.json()).rankings ?? [])
.map((r: { rwm_expected_pct: number | null }) => r.rwm_expected_pct)
.filter((v: number | null) => v != null);
expect(scores).toEqual([...scores].sort((a: number, b: number) => b - a));
});
/*
* Analytics on the location layer.
*
* Umami counts a pageview for every one of these URLs already. What it cannot
* say is which *kind* of location page earns engagement, because all four
* families share the /schools/ prefix — and that is the question that decides
* whether to keep investing in them.
*/
/** Capture Umami events, with the real script blocked so it cannot clobber
* the stub. Must be called before the first navigation. */
async function captureEvents(page: Page) {
const events: Array<{ name: string; data: Record<string, unknown> }> = [];
await page.route('**/analytics.schoolcompare.co.uk/**', (route) => route.abort());
await page.exposeFunction('__capture',
(name: string, data: Record<string, unknown>) => { events.push({ name, data }); });
await page.addInitScript(() => {
(window as unknown as { umami: unknown }).umami = {
track: (name: string, data: unknown) =>
(window as unknown as { __capture: (n: string, d: unknown) => void })
.__capture(name, data),
};
});
return events;
}
test('a location page reports which kind of place it is', async ({ page }) => {
const events = await captureEvents(page);
const place = await firstPlaceOfKind(page, 'authority');
await page.goto(`/schools/authority/${place.slug}`);
await expect.poll(() => events.find((e) => e.name === 'place_viewed'),
{ timeout: 10_000 }).toBeTruthy();
const event = events.find((e) => e.name === 'place_viewed')!;
expect(event.data.kind).toBe('authority');
expect(event.data.slug).toBe(place.slug);
expect(event.data.phase).toBe('all');
});
test('a school reached from a location page is attributed to it, not to direct', async ({ page }) => {
/*
* The defect this was written for. getNavigationSource had no case for
* /schools/, so every school view that came through the location layer was
* filed as 'direct' — the bucket you read as "typed the URL". The one
* measurement that says whether ~3,900 SEO pages work was reporting the
* wrong answer, confidently.
*/
const events = await captureEvents(page);
const place = await firstPlaceOfKind(page, 'town');
await page.goto(`/schools/${place.slug}`);
await page.locator('a[href^="/school/"]').first().click();
await page.waitForURL(/\/school\//);
await expect.poll(() => events.find((e) => e.name === 'school_viewed'),
{ timeout: 10_000 }).toBeTruthy();
expect(events.find((e) => e.name === 'school_viewed')!.data.from).toBe('place');
});
/*
* School autosuggest (spec 2026-08-26).
*/
async function autosuggestIsOn(page: Page): Promise<boolean> {
await page.goto('/');
return (await page.getByRole('combobox').count()) > 0;
}
test('the suggest endpoint answers from Typesense', async ({ page }) => {
// Not flagged — the endpoint is live even while the UI is dark, so it can
// be smoke-tested before the feature is switched on.
const res = await page.request.get('/api/suggest?q=brecknock');
expect(res.ok()).toBeTruthy();
const { suggestions } = await res.json();
expect(Array.isArray(suggestions)).toBeTruthy();
if (suggestions.length) {
// Local authority is what tells two "St Mary's" apart.
expect(suggestions[0]).toHaveProperty('school_name');
expect(suggestions[0]).toHaveProperty('local_authority');
}
});
test('a one-character query is answered, not rejected', async ({ page }) => {
// The keystroke path never errors on ordinary input.
const res = await page.request.get('/api/suggest?q=b');
expect(res.status()).toBe(200);
expect((await res.json()).suggestions).toEqual([]);
});
test('the suggest response is cacheable', async ({ page }) => {
const res = await page.request.get('/api/suggest?q=brecknock');
expect(res.headers()['cache-control'] ?? '').toContain('s-maxage');
});
test('typing a school name suggests it, and choosing it opens that school', async ({ page }) => {
test.skip(!(await autosuggestIsOn(page)),
'the school_autosuggest flag is off in this environment');
// A school certain to exist in any environment with data.
const { schools } = await (await page.request.get('/api/schools?page_size=1')).json();
test.skip(!schools?.length, 'no schools in this environment');
const name = schools[0].school_name as string;
await page.goto('/');
await page.getByRole('combobox').first().fill(name.slice(0, 12));
const option = page.getByRole('option').first();
await expect(option).toBeVisible();
await option.click();
await expect(page).toHaveURL(/\/school\/\d+/);
});
test('the whole dropdown is reachable, not clipped by the hero', async ({ page }) => {
/*
* The hero panel had overflow: hidden to clip its artwork to the rounded
* corners, and it clipped the dropdown too — 320px of list against 145px of
* panel below the input, so roughly half was cut off with nothing to say so.
*
* toBeVisible() does not catch this: it checks the box is non-empty and not
* visibility:hidden, and an ancestor's overflow clips neither. The invariant
* that does catch it is that the LAST option is the thing actually painted
* at its own coordinates — which fails for clipping and for occlusion alike.
*/
test.skip(!(await autosuggestIsOn(page)),
'the school_autosuggest flag is off in this environment');
const { schools } = await (await page.request.get('/api/schools?page_size=1')).json();
test.skip(!schools?.length, 'no schools in this environment');
await page.goto('/');
await page.getByRole('combobox').first().fill(
(schools[0].school_name as string).slice(0, 6));
const options = page.getByRole('option');
await expect(options.first()).toBeVisible();
const count = await options.count();
const painted = await options.nth(count - 1).evaluate((el) => {
const r = el.getBoundingClientRect();
const hit = document.elementFromPoint(r.left + r.width / 2, r.top + r.height / 2);
return { inside: el.contains(hit) || el === hit, bottom: Math.round(r.bottom) };
});
expect(painted.inside,
`the last option is not painted at its own coordinates (bottom ${painted.bottom}) `
+ '— an ancestor is clipping or covering the dropdown').toBeTruthy();
});
test('the dropdown does not survive into the results it produced', async ({ page }) => {
/*
* The bug that took the staging gate down, and it was not a test problem:
* after a search the results-page bar still holds the term, so the dropdown
* reopened on top of the results and swallowed the click on the first one.
* Playwright reported it as "<li role=option> intercepts pointer events"; a
* reader would simply have found their first result unclickable.
*/
test.skip(!(await autosuggestIsOn(page)),
'the school_autosuggest flag is off in this environment');
await page.goto('/');
await page.getByRole('combobox').first().fill('school');
await expect(page.getByRole('option').first()).toBeVisible();
await page.getByRole('button', { name: /Search/i }).first().click();
await page.waitForURL(/search=school/);
await expect(page.getByRole('listbox')).toHaveCount(0);
// And the results underneath are actually reachable, which is the point.
await page.locator('a[href^="/school/"]').first().click({ timeout: 15_000 });
await expect(page).toHaveURL(/\/school\//);
});
test('with autosuggest off, the search box is a plain input', async ({ page }) => {
test.skip(await autosuggestIsOn(page),
'the school_autosuggest flag is on in this environment');
await page.goto('/');
await expect(page.getByRole('combobox')).toHaveCount(0);
// And the box still works: the existing search must be untouched.
await page.getByPlaceholder(/School name or postcode/i).first().fill('abbey');
await page.getByRole('button', { name: /Search/i }).first().click();
await expect(page).toHaveURL(/search=abbey/);
});
// ── Destination measures ───────────────────────────────────────────────────
//
// Two failure modes have to be told apart here, and conflating them is how
// this suite would either hide a regression or block the promotion pipeline:
//
// * the backend does not serve the `destinations` field at all — a code
// regression, or a deploy that did not land. FAILS.
// * the field is served but every school is empty — the annual EES DAG has
// not run on this environment yet. SKIPS, loudly.
//
// The second is a data-load precondition, not a defect, and it is true for
// every commit between this merging and the DAG being triggered. Failing on it
// would redden the staging gate for unrelated work. This is not the quiet skip
// 4f01fbd removed from the distance journeys: that one hid a broken feature
// behind a flag check, whereas the assertion that the code is deployed and
// correctly shaped still runs here on every commit.
async function secondaryWithDestinations(page: Page): Promise<{
urn: string; destinations: any;
}> {
const res = await page.request.get('/api/schools?search=school&per_page=100');
expect(res.ok()).toBeTruthy();
const body = await res.json();
const urns: string[] = (body.schools ?? [])
.filter((s: { phase?: string; attainment_8_score?: number | null }) =>
s.phase === 'Secondary' && s.attainment_8_score != null)
.map((s: { urn: number }) => String(s.urn));
expect(urns.length).toBeGreaterThan(0);
let served = false;
for (const urn of urns.slice(0, 25)) {
const detail = await page.request.get(`/api/schools/${urn}`);
if (!detail.ok()) continue;
const data = await detail.json();
// The key must exist, even as null. Its absence means the backend in front
// of us does not know about destinations at all.
if ('destinations' in data) served = true;
if (data.destinations?.ks4) return { urn, destinations: data.destinations };
}
expect(served,
'GET /api/schools/{urn} served no `destinations` key at all — the backend '
+ 'is missing this feature, not merely missing its data').toBeTruthy();
test.skip(true,
'No school has destination data yet: the annual EES DAG has not run on '
+ 'this environment. The API shape is correct, so this is a data-load '
+ 'precondition rather than a regression.');
throw new Error('unreachable');
}
test('a secondary school page says where its Year 11 leavers went', async ({ page }) => {
const { urn } = await secondaryWithDestinations(page);
await page.goto(`/school/${urn}`);
const section = page.locator('#destinations');
await expect(section).toBeVisible({ timeout: 15_000 });
await expect(section.getByRole('heading', { name: 'After Year 11' })).toBeVisible();
// The section must date its own cohort: destinations run about two GCSE
// years behind the results above them, and an undated figure reads as stale.
await expect(section).toContainText(/20\d{2}\/\d{2}/);
});
test('the destinations bar is absent entirely whenever a figure is withheld', async ({ page }) => {
const { urn, destinations } = await secondaryWithDestinations(page);
await page.goto(`/school/${urn}`);
const section = page.locator('#destinations');
await expect(section).toBeVisible({ timeout: 15_000 });
const allGroup = destinations.ks4.groups.all;
const suppressed = (allGroup?.categories ?? [])
.filter((c: { status: string }) => c.status === 'suppressed');
if (suppressed.length > 0) {
// R1: a bar drawn from the published segments leaves a gap whose width is
// the withheld figure, readable straight off the axis.
await expect(section.locator('[data-destination-segment]')).toHaveCount(0);
await expect(section.getByText(/withheld/i).first()).toBeVisible();
} else {
const published = (allGroup?.categories ?? [])
.filter((c: { status: string }) => c.status === 'published');
await expect(section.locator('[data-destination-segment]'))
.toHaveCount(published.length);
}
});
test('switching to disadvantaged pupils never reveals a withheld figure', async ({ page }) => {
const { urn, destinations } = await secondaryWithDestinations(page);
const disadvantaged = destinations.ks4.groups.disadvantaged;
test.skip(!disadvantaged, 'this school publishes no disadvantaged breakdown');
await page.goto(`/school/${urn}`);
const section = page.locator('#destinations');
await expect(section).toBeVisible({ timeout: 15_000 });
const radio = section.getByRole('radio', { name: /disadvantaged/i });
await expect(radio).toBeVisible();
await radio.click();
const suppressed = (disadvantaged.categories ?? [])
.filter((c: { status: string }) => c.status === 'suppressed');
if (suppressed.length > 0) {
await expect(section.locator('[data-destination-segment]')).toHaveCount(0);
// The residual must appear nowhere on the page — it is the withheld figure.
const cohort: number = disadvantaged.cohort;
const publishedTotal = (disadvantaged.categories ?? [])
.filter((c: { status: string }) => c.status === 'published')
.reduce((sum: number, c: { pupils: number }) => sum + c.pupils, 0);
const residual = cohort - publishedTotal;
const text = (await section.textContent()) ?? '';
expect(text).not.toMatch(new RegExp(`\\b${residual}\\b`));
}
});
test('a school with no sixth form has no post-16 destinations section', async ({ page }) => {
const res = await page.request.get('/api/schools?search=school&per_page=100');
const body = await res.json();
const noSixthForm = (body.schools ?? [])
.filter((s: { phase?: string; has_sixth_form?: boolean }) =>
s.phase === 'Secondary' && s.has_sixth_form === false)
.map((s: { urn: number }) => String(s.urn));
test.skip(noSixthForm.length === 0, 'no sixth-form-less secondary in this dataset');
await page.goto(`/school/${noSixthForm[0]}`);
await expect(page.locator('h1').first()).toBeVisible({ timeout: 15_000 });
// Absence is the correct statement, so there must be no placeholder either.
await expect(page.locator('#post16-destinations')).toHaveCount(0);
await expect(page.getByText(/destination data coming soon/i)).toHaveCount(0);
});
test('the destinations section never claims a pupil stayed at this school', async ({ page }) => {
const { urn } = await secondaryWithDestinations(page);
await page.goto(`/school/${urn}`);
const section = page.locator('#destinations');
await expect(section).toBeVisible({ timeout: 15_000 });
// The published file records the TYPE of place a leaver went to, never which
// one, so the page can never say a pupil stayed on here.
const text = (await section.textContent()) ?? '';
expect(text).not.toMatch(/stayed on (here|at this school)/i);
});
/**
* The About page and the blog exist to give the site a named human author.
* These journeys assert the load-bearing parts of that — a name, a face, the
* honesty claim, and a resolvable Person entity — rather than exact copy,
* which will be edited.
*
* Both are behind flags (about_page, blog), so each has a lit journey and a
* dark one. Flag state is read from the observable effect rather than from
* /api/flags, which the public proxy denies on purpose — the same approach
* distanceFeatureIsOn() takes above.
*/
async function aboutPageIsOn(page: Page): Promise<boolean> {
return (await page.request.get('/about')).ok();
}
async function blogIsOn(page: Page): Promise<boolean> {
return (await page.request.get('/blog')).ok();
}
test('with the about page off, it is absent rather than empty', async ({ page }) => {
test.skip(await aboutPageIsOn(page), 'the about_page flag is on in this environment');
// Dark means the URL does not exist, not that it renders empty: a 404 is
// what stops a crawler keeping the page in its index.
expect((await page.request.get('/about')).status()).toBe(404);
// A footer link into a 404 is the failure this flag has to avoid.
await page.goto('/');
await expect(page.locator('footer a[href="/about"]')).toHaveCount(0);
// And a sitemap must never advertise a URL that 404s.
const sitemap = await page.request.get('/content-sitemap.xml');
expect(await sitemap.text()).not.toContain('/about');
});
test('with the blog off, it is absent rather than empty', async ({ page }) => {
test.skip(await blogIsOn(page), 'the blog flag is on in this environment');
expect((await page.request.get('/blog')).status()).toBe(404);
expect((await page.request.get('/blog/rss.xml')).status()).toBe(404);
await page.goto('/');
await expect(page.locator('footer a[href="/blog"]')).toHaveCount(0);
const sitemap = await page.request.get('/content-sitemap.xml');
expect(await sitemap.text()).not.toContain('/blog');
// The admin panel is deliberately NOT flagged: posts have to be writable
// before the blog is readable, or there is nothing to turn on.
expect((await page.request.get('/admin')).status()).not.toBe(404);
});
test('the about page names a human author and is reachable from the footer', async ({ page }) => {
test.skip(!(await aboutPageIsOn(page)), 'the about_page flag is off in this environment');
await page.goto('/');
const aboutLink = page.locator('footer a[href="/about"]');
await expect(aboutLink).toBeVisible();
await aboutLink.click();
await page.waitForURL(/\/about$/);
await expect(page.getByRole('heading', { level: 1 })).toContainText('Tudor');
await expect(page.locator('img[alt*="Tudor"]')).toBeVisible();
// The credibility claim is lived experience plus stated provenance, not
// expertise. If this sentence ever disappears the positioning has drifted.
await expect(page.getByText(/not an education expert/i)).toBeVisible();
const jsonLd = await page
.locator('script[type="application/ld+json"]')
.first()
.textContent();
expect(jsonLd).toContain('"Person"');
// First name only — a surname here would be the one place it leaks.
expect(jsonLd).not.toMatch(/familyName/);
});
test('the blog lists posts and each one renders with a byline', async ({ page }) => {
test.skip(!(await blogIsOn(page)), 'the blog flag is off in this environment');
await page.goto('/blog');
await expect(page.getByRole('heading', { level: 1 })).toBeVisible();
const postLinks = page.locator('a[href^="/blog/"]');
// Data invariant: staging must carry at least one published post. If this
// fails, the environment has no content rather than the code being broken.
expect(await postLinks.count()).toBeGreaterThan(0);
await postLinks.first().click();
await page.waitForURL(/\/blog\/.+/);
await expect(page.getByRole('heading', { level: 1 })).toBeVisible();
await expect(page.getByText(/^By Tudor/)).toBeVisible();
const jsonLd = await page
.locator('script[type="application/ld+json"]')
.first()
.textContent();
expect(jsonLd).toContain('"BlogPosting"');
});
test('the admin panel is not indexable', async ({ page }) => {
const response = await page.request.get('/admin');
expect(response.headers()['x-robots-tag']).toContain('noindex');
});
test('the content sitemap lists the about page and is advertised in robots', async ({ page }) => {
const sitemap = await page.request.get('/content-sitemap.xml');
// Served whatever the flags say: robots.txt names it unconditionally, and
// with both dark it is a valid empty urlset rather than a 404.
expect(sitemap.ok()).toBeTruthy();
if (await aboutPageIsOn(page)) {
expect(await sitemap.text()).toContain('/about');
}
// The school corpus sitemap is proxied from FastAPI; this one is Next's.
// robots.txt must advertise both or the blog never gets discovered.
const robots = await page.request.get('/robots.txt');
const body = await robots.text();
expect(body).toContain('/sitemap.xml');
expect(body).toContain('/content-sitemap.xml');
});
+1
View File
@@ -39,3 +39,4 @@ yarn-error.log*
# typescript # typescript
*.tsbuildinfo *.tsbuildinfo
next-env.d.ts next-env.d.ts
+7
View File
@@ -53,6 +53,13 @@ COPY --from=builder /app/.next/static ./.next/static
# a miss here is a silent 500 on /opengraph-image, not a build failure. # a miss here is a silent 500 on /opengraph-image, not a build failure.
COPY --from=builder /app/assets ./assets COPY --from=builder /app/assets ./assets
# Payload writes uploads here, and the compose file mounts a named volume over
# it. The directory must exist and be owned by the runtime user BEFORE the
# mount: Docker seeds a fresh named volume from the image path, so a missing or
# root-owned directory here makes every upload fail with EACCES at runtime,
# long after the build passed. The chown below covers it.
RUN mkdir -p /app/media
# Set correct permissions # Set correct permissions
RUN chown -R nextjs:nodejs /app RUN chown -R nextjs:nodejs /app
@@ -0,0 +1,31 @@
/**
* The /api/* proxy is public. Anything it forwards is on the internet.
*
* @jest-environment node
*/
// The docblock above is load-bearing. jest.config.js sets jsdom globally, and
// NextRequest/NextResponse need the Web Fetch API globals that only the node
// environment provides — under jsdom this suite fails on import, not on an
// assertion.
import { NextRequest } from 'next/server';
import { GET } from '@/app/(frontend)/api/[...path]/route';
function request(path: string) {
return new NextRequest(`http://localhost:3000/api/${path}`);
}
describe('public API proxy', () => {
it('refuses to forward internal-only paths', async () => {
// /api/flags names every unreleased feature and its state. Forwarding it
// publishes the thing shipping dark exists to keep quiet.
const res = await GET(request('flags'), { params: Promise.resolve({ path: ['flags'] }) });
expect(res.status).toBe(404);
});
it('does not deny a path that merely starts with the same letters', async () => {
// A prefix match would take /api/flagship down with /api/flags.
const res = await GET(
request('flagship'), { params: Promise.resolve({ path: ['flagship'] }) });
expect(res.status).not.toBe(404);
});
});
@@ -0,0 +1,34 @@
import { metadata } from '@/app/(frontend)/about/page';
import { personJsonLd, organizationJsonLd } from '@/lib/jsonld';
describe('/about metadata', () => {
it('canonicalises to the bare path', () => {
expect(metadata.alternates?.canonical)
.toBe('https://www.schoolcompare.co.uk/about');
});
});
describe('author structured data', () => {
it('describes a Person with a first name and a photo', () => {
const person = personJsonLd();
expect(person['@type']).toBe('Person');
expect(person.name).toBe('Tudor');
expect(person.image).toBe('https://www.schoolcompare.co.uk/brand/tudor.jpg');
expect(person.url).toBe('https://www.schoolcompare.co.uk/about');
});
it('never publishes a surname or an employer', () => {
// Author identity constraint: first name only. A surname here would be
// the one place it leaks, since JSON-LD is machine-read and archived.
const serialised = JSON.stringify(personJsonLd());
expect(serialised).not.toMatch(/familyName|Sitaru/i);
expect(serialised).not.toMatch(/worksFor|affiliation/i);
});
it('describes the site as an Organization the Person authors for', () => {
const org = organizationJsonLd();
expect(org['@type']).toBe('Organization');
expect(org.name).toBe('schoolcompare');
expect(org.url).toBe('https://www.schoolcompare.co.uk');
});
});
@@ -0,0 +1,66 @@
/**
* The blog index imports getCachedPayload, which pulls in Payload — ESM-only,
* and next/jest will not transform node_modules. Mocking that one module keeps
* the page's metadata testable without loading the CMS; the mock is never
* called, because `metadata` is a static export evaluated at import time.
*/
jest.mock('@/lib/payload', () => ({ getCachedPayload: jest.fn() }));
import { metadata } from '@/app/(frontend)/blog/page';
import { blogPostingJsonLd, breadcrumbJsonLd } from '@/lib/jsonld';
const post = {
title: 'What the data cannot tell you',
slug: 'what-the-data-cannot-tell-you',
excerpt: 'Results describe one year group on a handful of days.',
publishedAt: '2026-09-15T00:00:00.000Z',
};
describe('/blog metadata', () => {
it('canonicalises to the bare path', () => {
expect(metadata.alternates?.canonical)
.toBe('https://www.schoolcompare.co.uk/blog');
});
});
describe('BlogPosting structured data', () => {
it('names the same Person entity the about page declares', () => {
// By @id, not by repeating the person: search engines must resolve every
// post and the about page to one author entity, or the site has several.
const ld = blogPostingJsonLd(post, { namedAuthor: true });
expect(ld['@type']).toBe('BlogPosting');
expect(ld.author['@id']).toBe('https://www.schoolcompare.co.uk/about#tudor');
expect(ld.publisher['@id']).toBe('https://www.schoolcompare.co.uk#organization');
});
it('attributes to the organization when the about page is dark', () => {
/*
* The two flags are independent, so blog-on-about-off is a reachable
* state. The Person entity lives at /about#tudor and that URL 404s while
* the flag is dark, so claiming it would declare an author that resolves
* to nothing — worse for the blog's credibility than having no named
* author at all. Attribute to the publisher instead.
*/
const ld = blogPostingJsonLd(post, { namedAuthor: false });
expect(ld.author['@id']).toBe('https://www.schoolcompare.co.uk#organization');
expect(JSON.stringify(ld)).not.toContain('/about');
});
it('carries a self-referencing canonical url and the publish date', () => {
const ld = blogPostingJsonLd(post, { namedAuthor: true });
expect(ld.url).toBe(
'https://www.schoolcompare.co.uk/blog/what-the-data-cannot-tell-you',
);
expect(ld.datePublished).toBe('2026-09-15T00:00:00.000Z');
});
});
describe('breadcrumbs', () => {
it('places the post under the blog index', () => {
const ld = breadcrumbJsonLd(post);
expect(ld.itemListElement[0].item).toBe('https://www.schoolcompare.co.uk/blog');
expect(ld.itemListElement[1].item).toBe(
'https://www.schoolcompare.co.uk/blog/what-the-data-cannot-tell-you',
);
});
});
+4 -4
View File
@@ -1,7 +1,7 @@
import { metadata as homeMetadata } from '@/app/page'; import { metadata as homeMetadata } from '@/app/(frontend)/page';
import { metadata as rankingsMetadata } from '@/app/rankings/page'; import { metadata as rankingsMetadata } from '@/app/(frontend)/rankings/page';
import { metadata as admissionsMetadata } from '@/app/admissions/page'; import { metadata as admissionsMetadata } from '@/app/(frontend)/admissions/page';
import { generateMetadata as compareMetadata } from '@/app/compare/page'; import { generateMetadata as compareMetadata } from '@/app/(frontend)/compare/page';
describe('canonical URLs', () => { describe('canonical URLs', () => {
it('the homepage canonicalises to the bare root', () => { it('the homepage canonicalises to the bare root', () => {
@@ -0,0 +1,66 @@
/**
* next.config.mjs carries the staging noindex rule. Breaking it turns
* stx.schoolcompare.co.uk into a fully crawlable duplicate of production,
* and nothing else in the suite would notice.
*
* The non-null assertions are deliberate: every key asserted here is optional
* on NextConfig, and a missing one is precisely the regression under test, so
* the assertion below should fail the test rather than the compile.
*/
import nextConfig from '@/next.config.mjs';
async function headerRules() {
return nextConfig.headers!();
}
describe('next.config.mjs', () => {
it('keeps the staging host out of the index', async () => {
const headers = await headerRules();
const stagingRule = headers.find((rule) =>
rule.has?.some(
(cond) => cond.type === 'host' && cond.value === 'stx.schoolcompare.co.uk',
),
);
expect(stagingRule).toBeDefined();
expect(stagingRule!.headers).toContainEqual({
key: 'X-Robots-Tag',
value: 'noindex, nofollow',
});
});
it('still emits standalone output for the Docker runner', () => {
expect(nextConfig.output).toBe('standalone');
});
it('still traces the share-card fonts into the standalone bundle', () => {
expect(nextConfig.outputFileTracingIncludes!['/opengraph-image']).toEqual([
'./assets/**',
]);
});
it('still allows the analytics subdomain to frame the site', async () => {
const headers = await headerRules();
const csp = headers
.flatMap((rule) => rule.headers)
.find((header) => header.key === 'Content-Security-Policy');
expect(csp).toBeDefined();
expect(csp!.value).toContain('https://analytics.schoolcompare.co.uk');
});
});
describe('admin surface', () => {
it('serves noindex on the admin panel and the CMS API', async () => {
// robots.txt disallows these too, but a Disallow only blocks crawling — a
// URL found from an external link can still be indexed without ever being
// fetched. This header is what actually keeps them out.
const headers = await headerRules();
for (const source of ['/admin/:path*', '/cms-api/:path*']) {
const rule = headers.find((entry) => entry.source === source);
expect(rule).toBeDefined();
expect(rule!.headers).toContainEqual({
key: 'X-Robots-Tag',
value: 'noindex, nofollow',
});
}
});
});
@@ -1,4 +1,4 @@
import { generateMetadata as placeMeta } from '@/app/schools/[place]/page'; import { generateMetadata as placeMeta } from '@/app/(frontend)/schools/[place]/page';
jest.mock('@/lib/places', () => ({ jest.mock('@/lib/places', () => ({
...jest.requireActual('@/lib/places'), ...jest.requireActual('@/lib/places'),
+20
View File
@@ -0,0 +1,20 @@
import robots from '@/app/robots';
describe('robots.txt', () => {
it('disallows the admin panel and the CMS API', () => {
const rules = robots().rules;
const rule = Array.isArray(rules) ? rules[0] : rules;
expect(rule.disallow).toEqual(
expect.arrayContaining(['/api/', '/_next/', '/admin/', '/cms-api/']),
);
});
});
describe('sitemap discovery', () => {
it('lists both the proxied school sitemap and the Next-owned content sitemap', () => {
expect(robots().sitemap).toEqual([
'https://www.schoolcompare.co.uk/sitemap.xml',
'https://www.schoolcompare.co.uk/content-sitemap.xml',
]);
});
});
@@ -0,0 +1,149 @@
import { render, screen } from '@testing-library/react';
import { DestinationsSection } from '@/components/school/DestinationsSection';
import type { DestinationPhase } from '@/lib/types';
import type { DestinationCategory, DestinationStatus } from '@/lib/destinations';
const cell = (
category: DestinationCategory,
pupils: number | null,
status: DestinationStatus = 'published',
) => ({
category, pupils,
percentage: pupils === null ? null : (pupils / 180) * 100,
status,
});
const ALL_PUBLISHED = [
cell('school_sixth_form', 75), cell('sixth_form_college', 21),
cell('further_education', 55), cell('other_education', 6),
cell('apprenticeship', 8), cell('employment', 6),
cell('not_sustained', 5), cell('not_captured', 4),
];
const fullPhase: DestinationPhase = {
cohort_year: '2022/23',
groups: { all: { cohort: 180, categories: ALL_PUBLISHED } },
};
const suppressedPhase: DestinationPhase = {
cohort_year: '2022/23',
groups: {
all: {
cohort: 180,
categories: [
cell('school_sixth_form', 75), cell('sixth_form_college', null, 'suppressed'),
cell('further_education', 55), cell('other_education', 6),
cell('apprenticeship', 8), cell('employment', 6),
cell('not_sustained', 5), cell('not_captured', 4),
],
},
},
};
describe('DestinationsSection', () => {
it('dates its own cohort so it is not read as stale next to the GCSE section', () => {
render(<DestinationsSection destinations={fullPhase} />);
expect(screen.getByText(/2022\/23/)).toBeInTheDocument();
});
it('renders one bar segment per published category', () => {
const { container } = render(<DestinationsSection destinations={fullPhase} />);
expect(container.querySelectorAll('[data-destination-segment]')).toHaveLength(8);
});
it('renders NO bar at all when a category is withheld', () => {
const { container } = render(<DestinationsSection destinations={suppressedPhase} />);
// R1: a bar with a gap in it publishes the withheld figure by its width.
expect(container.querySelectorAll('[data-destination-segment]')).toHaveLength(0);
expect(screen.getAllByText(/withheld/i).length).toBeGreaterThan(0);
});
it('never states the remainder for a partially suppressed group', () => {
const { container } = render(<DestinationsSection destinations={suppressedPhase} />);
// 180 cohort - 159 published = 21, the withheld figure. It must appear nowhere.
expect(container.textContent).not.toMatch(/\b21\b/);
});
it('shows a card value for a group whose components are all published', () => {
render(<DestinationsSection destinations={fullPhase} />);
// academic route = 75 + 21 = 96 of 180 = 53%
expect(screen.getByText('53%')).toBeInTheDocument();
});
it('refuses a card value when one of its components is withheld', () => {
render(<DestinationsSection destinations={suppressedPhase} />);
// academic route needs sixth_form_college, which is suppressed.
expect(screen.getByText(/not published/i)).toBeInTheDocument();
expect(screen.queryByText('53%')).not.toBeInTheDocument();
});
it('never claims a pupil stayed at this school', () => {
const { container } = render(<DestinationsSection destinations={fullPhase} />);
// The published file reports destination TYPE, never destination institution.
expect(container.textContent).not.toMatch(/stayed on (here|at this school)/i);
});
it('renders nothing when no group carries categories', () => {
const empty: DestinationPhase = { cohort_year: '2022/23', groups: {} };
const { container } = render(<DestinationsSection destinations={empty} />);
expect(container.firstChild).toBeNull();
});
});
describe('the detail table keeps the three statuses apart', () => {
// 'suppressed' and 'not_applicable' are different claims, and the mart, the
// SQLAlchemy model and the serialiser all preserve the difference. The table
// used to key its Share column off `percentage === null`, which is true for
// both, so a category that simply does not apply was labelled "withheld" —
// while the Pupils column beside it rendered blank.
const mixedPhase: DestinationPhase = {
cohort_year: '2022/23',
groups: {
all: {
cohort: 180,
categories: [
cell('school_sixth_form', 75),
cell('sixth_form_college', null, 'suppressed'),
cell('further_education', null, 'suppressed'),
cell('apprenticeship', null, 'not_applicable'),
],
},
},
};
const rowFor = (container: HTMLElement, category: string) =>
Array.from(container.querySelectorAll('tbody tr'))
.find(tr => tr.textContent?.includes(category));
it('never labels a not-applicable category as withheld', () => {
const { container } = render(<DestinationsSection destinations={mixedPhase} />);
const row = rowFor(container, 'Apprenticeship');
expect(row).toBeTruthy();
expect(row!.textContent).not.toMatch(/withheld/i);
});
it('labels a genuinely suppressed category as withheld in both columns', () => {
const { container } = render(<DestinationsSection destinations={mixedPhase} />);
const row = rowFor(container, 'Sixth-form college');
expect(row).toBeTruthy();
expect(row!.querySelectorAll('td')).toHaveLength(2);
Array.from(row!.querySelectorAll('td')).forEach(td =>
expect(td.textContent).toMatch(/withheld/i));
});
it('the two columns of a row never disagree about what the row is', () => {
const { container } = render(<DestinationsSection destinations={mixedPhase} />);
Array.from(container.querySelectorAll('tbody tr')).forEach(tr => {
const cells = Array.from(tr.querySelectorAll('td'))
.map(td => /withheld/i.test(td.textContent ?? ''));
expect(new Set(cells).size).toBe(1);
});
});
it('shows a published category its real figures', () => {
const { container } = render(<DestinationsSection destinations={mixedPhase} />);
const row = rowFor(container, 'State-funded school sixth form');
expect(row!.textContent).toMatch(/75/);
expect(row!.textContent).toMatch(/42%/);
});
});
@@ -0,0 +1,110 @@
import { render, screen, fireEvent, waitFor } from '@testing-library/react';
import userEvent from '@testing-library/user-event';
import { FilterBar } from '@/components/FilterBar';
const push = jest.fn();
let searchParams = new URLSearchParams();
jest.mock('next/navigation', () => ({
useRouter: () => ({ push, replace: jest.fn(), prefetch: jest.fn() }),
usePathname: () => '/',
useSearchParams: () => searchParams,
}));
const FILTERS = {
local_authorities: [], school_types: [], years: [], phases: [],
genders: [], admissions_policies: [],
};
const realFetch = global.fetch;
beforeEach(() => {
global.fetch = jest.fn(async () => ({
ok: true,
json: async () => ({ suggestions: [{
urn: 100010, school_name: 'Brecknock Primary School',
local_authority: 'Camden', postcode: 'NW1 1AA',
phase: 'Primary', school_type: 'Community school' }] }),
})) as unknown as typeof fetch;
push.mockClear();
searchParams = new URLSearchParams();
});
afterEach(() => { global.fetch = realFetch; });
describe('FilterBar autosuggest', () => {
it('is a combobox only when the flag is on', () => {
const { rerender } = render(<FilterBar filters={FILTERS} autosuggest={false} />);
expect(screen.queryByRole('combobox')).not.toBeInTheDocument();
rerender(<FilterBar filters={FILTERS} autosuggest />);
expect(screen.getByRole('combobox')).toBeInTheDocument();
});
it('makes no request while the flag is off', async () => {
// Off means off: no listener, no fetch, no markup.
render(<FilterBar filters={FILTERS} autosuggest={false} />);
await userEvent.type(screen.getByPlaceholderText(/School name or postcode/i),
'brecknock');
expect(global.fetch).not.toHaveBeenCalled();
});
it('shows suggestions and navigates when one is chosen', async () => {
render(<FilterBar filters={FILTERS} autosuggest />);
await userEvent.type(screen.getByRole('combobox'), 'brecknock');
const option = await screen.findByRole('option', { name: /Brecknock/ });
await userEvent.click(option);
expect(push).toHaveBeenCalledWith(
expect.stringContaining('/school/100010'));
});
it('suppresses suggestions once the value is a postcode', async () => {
// The box takes a name OR a postcode; suggestions must get out of the way.
//
// fireEvent.change, not userEvent.type: typing sets "N", "NW", "NW1"... and
// "NW1" is not a postcode, so a request for it is correct behaviour. Only
// the settled value is the assertion, so set it in one go.
render(<FilterBar filters={FILTERS} autosuggest />);
fireEvent.change(screen.getByRole('combobox'), { target: { value: 'NW1 1AA' } });
await new Promise((r) => setTimeout(r, 300)); // past the 200ms debounce
expect(global.fetch).not.toHaveBeenCalled();
});
it('Enter with no active option still submits the free-text search', async () => {
// The existing behaviour is preserved, not replaced.
render(<FilterBar filters={FILTERS} autosuggest />);
const input = screen.getByRole('combobox');
await userEvent.type(input, 'brecknock{Enter}');
// updateURL pushes inside startTransition, so the call is not synchronous.
await waitFor(() => expect(push).toHaveBeenCalledWith(
expect.stringContaining('search=brecknock')));
});
});
describe('FilterBar autosuggest does not reopen over results', () => {
it('stays shut when the input arrives pre-filled from the URL', async () => {
/*
* The results-page bar renders with the search term already in the input.
* Opening on that would drop the dropdown on top of the results the search
* just produced — which is exactly what happened: the first result became
* unclickable, because the list sat over it and swallowed the pointer.
*
* Suggestions answer typing, not the presence of a value.
*/
searchParams = new URLSearchParams('search=brecknock');
render(<FilterBar filters={FILTERS} autosuggest />);
expect(screen.getByRole('combobox')).toHaveValue('brecknock');
await new Promise((r) => setTimeout(r, 300)); // past the 200ms debounce
expect(global.fetch).not.toHaveBeenCalled();
expect(screen.queryByRole('listbox')).not.toBeInTheDocument();
});
it('closes the dropdown when the search is submitted', async () => {
render(<FilterBar filters={FILTERS} autosuggest />);
const input = screen.getByRole('combobox');
await userEvent.type(input, 'brecknock');
expect(await screen.findByRole('listbox')).toBeInTheDocument();
await userEvent.type(input, '{Enter}');
await waitFor(() =>
expect(screen.queryByRole('listbox')).not.toBeInTheDocument());
});
});
@@ -0,0 +1,47 @@
/**
* The footer is the only navigational route to /about and /blog, so it is
* where a dark flag would otherwise leave a link into a 404.
*
* Both props default to false. A caller that forgets to pass them hides the
* links, which is the direction that cannot break a page — the same reasoning
* as backend/flags.py's "every flag defaults to False".
*/
import { render, screen } from '@testing-library/react';
import { Footer } from '@/components/Footer';
describe('footer feature links', () => {
it('links to both when both flags are on', () => {
render(<Footer aboutEnabled blogEnabled />);
expect(screen.getByRole('link', { name: /who's behind this/i }))
.toHaveAttribute('href', '/about');
expect(screen.getByRole('link', { name: /^blog$/i }))
.toHaveAttribute('href', '/blog');
});
it('omits the about link when that flag is dark', () => {
render(<Footer blogEnabled />);
expect(screen.queryByRole('link', { name: /who's behind this/i })).toBeNull();
expect(screen.getByRole('link', { name: /^blog$/i })).toBeInTheDocument();
});
it('omits the blog link when that flag is dark', () => {
render(<Footer aboutEnabled />);
expect(screen.queryByRole('link', { name: /^blog$/i })).toBeNull();
expect(screen.getByRole('link', { name: /who's behind this/i }))
.toBeInTheDocument();
});
it('drops the whole section when both are dark, not an empty heading', () => {
// Shipping dark means the footer renders as it did before the feature
// existed, not as a section with its contents removed.
render(<Footer />);
expect(screen.queryByRole('heading', { name: /^about$/i })).toBeNull();
expect(screen.queryByRole('link', { name: /who's behind this/i })).toBeNull();
expect(screen.queryByRole('link', { name: /^blog$/i })).toBeNull();
});
it('defaults to dark when a caller passes nothing', () => {
render(<Footer />);
expect(screen.queryByRole('link', { name: /who's behind this/i })).toBeNull();
});
});
@@ -199,3 +199,281 @@ describe('PlaceView table alignment', () => {
expect(container.querySelectorAll('th')[0].className).toBe(''); expect(container.querySelectorAll('th')[0].className).toBe('');
}); });
}); });
describe('PlaceView authorities', () => {
const straddling: PlaceDetail = {
place: { kind: 'outcode', slug: 'sw19', name: 'SW19', count: 33,
parent_authority: 'Merton', phases: ['primary'],
authorities: [
{ name: 'Merton', slug: 'merton', count: 26 },
{ name: 'Wandsworth', slug: 'wandsworth', count: 7 },
] },
schools: [
{ urn: 1, school_name: 'Alpha Primary', phase: 'Primary',
rwm_expected_pct: 82, attainment_8_score: null } as never,
],
averages: { rwm_expected_pct: 63, attainment_8_score: null },
};
it('names every authority the place straddles, not just the largest', () => {
// SW19 is mostly Merton but partly Wandsworth. Naming one asserts
// something false about a quarter of outcodes.
render(<PlaceView detail={straddling} englandAverage={61} neighbours={[]} />);
expect(screen.getByRole('link', { name: 'Merton' }))
.toHaveAttribute('href', '/schools/authority/merton');
expect(screen.getByRole('link', { name: 'Wandsworth' }))
.toHaveAttribute('href', '/schools/authority/wandsworth');
});
it('joins them readably rather than as a bare list', () => {
// Asserted on the summary line's whole text: a loose /and/ matcher also
// hits "Wandsworth".
const { container } = render(<PlaceView detail={straddling}
englandAverage={61} neighbours={[]} />);
const summary = container.querySelector('header p');
expect(summary?.textContent).toContain('Merton and Wandsworth');
});
it('falls back to the single parent when the field is absent', () => {
// A cached API response predating the authorities field must not blank
// the line entirely.
const legacy = { ...straddling,
place: { ...straddling.place, authorities: undefined } };
render(<PlaceView detail={legacy} englandAverage={61} neighbours={[]} />);
expect(screen.getByRole('link', { name: 'Merton' })).toBeInTheDocument();
});
});
describe('PlaceView list ordering', () => {
const detail3: PlaceDetail = {
place: { kind: 'town', slug: 'brentwood', name: 'Brentwood', count: 2,
parent_authority: 'Essex', phases: ['primary'] },
schools: [
{ urn: 1, school_name: 'Alpha Primary', phase: 'Primary',
rwm_expected_pct: 40, attainment_8_score: null } as never,
{ urn: 2, school_name: 'Beta Primary', phase: 'Primary',
rwm_expected_pct: 90, attainment_8_score: null } as never,
],
averages: { rwm_expected_pct: 65, attainment_8_score: null },
};
it('renders schools in the order the API sent them, not by score', () => {
// The API sorts alphabetically now; the component must not re-sort.
render(<PlaceView detail={detail3} englandAverage={61} neighbours={[]} />);
const links = screen.getAllByRole('link', { name: /Primary$/ });
expect(links.map((l) => l.textContent))
.toEqual(['Alpha Primary', 'Beta Primary']);
});
it('declares the list as ascending rather than implying a ranking', () => {
// An ItemList carrying `position` reads as a ranking unless it says
// otherwise, and the table is A-Z.
const { container } = render(<PlaceView detail={detail3} englandAverage={61}
neighbours={[]} />);
const ld = JSON.parse(
container.querySelector('script[type="application/ld+json"]')!.textContent!);
const list = ld['@graph'].find((n: { '@type': string }) => n['@type'] === 'ItemList');
expect(list.itemListOrder).toBe('https://schema.org/ItemListOrderAscending');
});
});
describe('PlaceView phase links', () => {
const authority: PlaceDetail = {
place: { kind: 'authority', slug: 'barnet', name: 'Barnet', count: 156,
parent_authority: null, phases: ['primary', 'secondary'] },
schools: [
{ urn: 1, school_name: 'Alpha Primary', phase: 'Primary',
rwm_expected_pct: 82, attainment_8_score: null } as never,
],
averages: { rwm_expected_pct: 63, attainment_8_score: null },
};
it('keeps an authority phase link in the authority namespace', () => {
// The link was built as `/schools/${slug}/${phase}` for every kind, so an
// authority page pointed into the town namespace. For 87 of 151
// authorities that 404'd; for the other 64 it silently landed on the town
// page of the same name — a different set of schools, and exactly the
// duplicate the two namespaces exist to prevent. Barnet is one of the 64.
render(<PlaceView detail={authority} englandAverage={61} neighbours={[]} />);
expect(screen.getByRole('link', { name: /^Primary schools in Barnet$/ }))
.toHaveAttribute('href', '/schools/authority/barnet/primary');
expect(screen.getByRole('link', { name: /^Secondary schools in Barnet$/ }))
.toHaveAttribute('href', '/schools/authority/barnet/secondary');
});
it('still uses the bare namespace for a town', () => {
render(<PlaceView detail={detail} englandAverage={61} neighbours={[]} />);
expect(screen.getByRole('link', { name: /^Primary schools in Brentwood$/ }))
.toHaveAttribute('href', '/schools/brentwood/primary');
});
it('offers no phase link when the place publishes none', () => {
// Outcodes are the case: no phase route exists for them, so the registry
// reports no phases and the nav does not render.
const outcode = { ...detail,
place: { ...detail.place, kind: 'outcode', slug: 'cm13', name: 'CM13',
phases: [] } };
render(<PlaceView detail={outcode} englandAverage={61} neighbours={[]} />);
expect(screen.queryByRole('navigation', { name: 'By phase' }))
.not.toBeInTheDocument();
});
});
describe('PlaceView unlinkable authorities', () => {
const withUnpublished: PlaceDetail = {
place: { kind: 'outcode', slug: 'tr21', name: 'TR21', count: 8,
parent_authority: 'Cornwall', phases: [],
authorities: [
{ name: 'Cornwall', slug: 'cornwall', count: 6 },
{ name: 'Isles Of Scilly', slug: null, count: 2 },
] },
schools: [
{ urn: 1, school_name: 'Alpha Primary', phase: 'Primary',
rwm_expected_pct: 82, attainment_8_score: null } as never,
],
averages: { rwm_expected_pct: 63, attainment_8_score: null },
};
it('names an authority with no page without linking it', () => {
// City of London and the Isles of Scilly hold fewer schools than a page
// needs. Saying where the place is stays right; linking there would 404.
const { container } = render(<PlaceView detail={withUnpublished}
englandAverage={61} neighbours={[]} />);
expect(screen.getByRole('link', { name: 'Cornwall' })).toBeInTheDocument();
expect(screen.queryByRole('link', { name: 'Isles Of Scilly' }))
.not.toBeInTheDocument();
expect(container.querySelector('header p')?.textContent)
.toContain('Isles Of Scilly');
});
});
describe('PlaceView school attributes', () => {
/*
* The table shipped with one column of scores, which answers "how did they
* do" and nothing about whether the school is one a family could use. Age
* range, faith, nursery and constituency are the four facts a parent
* filters on before they look at a number at all.
*/
const withAttributes: PlaceDetail = {
place: { kind: 'town', slug: 'chelmsford', name: 'Chelmsford', count: 3,
parent_authority: 'Essex', phases: ['primary', 'secondary'] },
schools: [
{ urn: 1, school_name: 'Alpha Primary', phase: 'Primary',
rwm_expected_pct: 82, attainment_8_score: null,
age_range: '4-11', religious_denomination: 'Church of England',
nursery_provision: true,
parliamentary_constituency: 'Chelmsford' } as never,
{ urn: 2, school_name: 'Beta High', phase: 'Secondary',
rwm_expected_pct: null, attainment_8_score: 47,
age_range: '11-16', religious_denomination: 'Does not apply',
nursery_provision: false,
parliamentary_constituency: 'Witham' } as never,
],
averages: { rwm_expected_pct: 63, attainment_8_score: 45 },
};
function headings(container: HTMLElement, table = 0): string[] {
return Array.from(container.querySelectorAll('table')[table]
.querySelectorAll('thead th')).map((th) => th.textContent ?? '');
}
it('heads a primary table with all four attributes', () => {
const { container } = render(<PlaceView detail={withAttributes}
englandAverage={61} neighbours={[]} />);
expect(headings(container)).toEqual([
'School', 'Reading, writing & maths',
'Ages', 'Religious character', 'Nursery', 'Constituency',
]);
});
it('omits nursery from a secondary table, where it does not apply', () => {
const { container } = render(<PlaceView detail={withAttributes}
englandAverage={61} neighbours={[]} />);
expect(headings(container, 1)).toEqual([
'School', 'Attainment 8', 'Ages', 'Religious character', 'Constituency',
]);
});
it('keeps the measure beside the school name, where a phone can see it', () => {
// Six columns overflow a phone and .tableWrap turns that into a swipe.
// With the measure last, the one number the page exists for is the one
// scrolled off the screen.
const { container } = render(<PlaceView detail={withAttributes}
phase="primary" englandAverage={61} neighbours={[]} />);
expect(headings(container)[1]).toBe('Reading, writing & maths');
});
it('shows the age range without repeating the column heading', () => {
render(<PlaceView detail={withAttributes} englandAverage={61}
neighbours={[]} />);
expect(screen.getByText('4–11')).toBeInTheDocument();
expect(screen.queryByText('Ages 4–11')).not.toBeInTheDocument();
});
it('names the faith of a faith school', () => {
render(<PlaceView detail={withAttributes} englandAverage={61}
neighbours={[]} />);
expect(screen.getByText('Church of England')).toBeInTheDocument();
});
it('reads "Does not apply" as no religious character, not as a value', () => {
// GIAS spells the absence of a faith as "Does not apply", which is a
// database answer rather than an English one. The school page already
// suppresses it; the two must not disagree about the same school.
const { container } = render(<PlaceView detail={withAttributes}
englandAverage={61} neighbours={[]} />);
const secondary = container.querySelectorAll('table')[1]
.querySelectorAll('tbody td');
expect(secondary[3].textContent).toBe('—');
expect(screen.queryByText(/Does not apply/)).not.toBeInTheDocument();
});
it('marks a nursery as such and a school without one as not', () => {
const { container } = render(<PlaceView detail={withAttributes}
englandAverage={61} neighbours={[]} />);
const cells = container.querySelectorAll('table')[0]
.querySelectorAll('tbody td');
expect(cells[4].textContent).toBe('Yes');
});
it('names the constituency of each school', () => {
render(<PlaceView detail={withAttributes} englandAverage={61}
neighbours={[]} />);
expect(screen.getByText('Chelmsford', { selector: 'td' })).toBeInTheDocument();
expect(screen.getByText('Witham', { selector: 'td' })).toBeInTheDocument();
});
it('dashes an attribute the data does not carry', () => {
// nursery_provision and parliamentary_constituency are absent from marts
// the pipeline has not rebuilt, and the API degrades them to null rather
// than failing. A row must survive that.
const bare: PlaceDetail = {
...withAttributes,
schools: [{ urn: 3, school_name: 'Gamma Primary', phase: 'Primary',
rwm_expected_pct: 70 } as never],
};
const { container } = render(<PlaceView detail={bare} phase="primary"
englandAverage={61} neighbours={[]} />);
const cells = Array.from(container.querySelectorAll('tbody td'))
.map((td) => td.textContent);
expect(cells.slice(2)).toEqual(['—', '—', '—', '—']);
});
it('gives an all-through school its nursery under primary only', () => {
// All-through schools render in both groups. Nursery belongs to the
// primary reading of the same school, not the secondary one.
const allThrough: PlaceDetail = {
...withAttributes,
schools: [{ urn: 4, school_name: 'Delta Academy', phase: 'All-through',
rwm_expected_pct: 66, attainment_8_score: 51,
age_range: '4-18', religious_denomination: 'None',
nursery_provision: true,
parliamentary_constituency: 'Chelmsford' } as never],
};
const { container } = render(<PlaceView detail={allThrough}
englandAverage={61} neighbours={[]} />);
const tables = container.querySelectorAll('table');
expect(tables[0].textContent).toContain('Yes');
expect(tables[1].textContent).not.toContain('Yes');
});
});
@@ -0,0 +1,44 @@
import { render, screen } from '@testing-library/react';
import { Post16DestinationsSection } from '@/components/school/Post16DestinationsSection';
import type { DestinationPhase } from '@/lib/types';
const phase: DestinationPhase = {
cohort_year: '2022/23',
groups: {
all: {
cohort: 96,
categories: [
{ category: 'higher_education', pupils: 56, percentage: 58.3, status: 'published' },
{ category: 'further_education', pupils: 12, percentage: 12.5, status: 'published' },
{ category: 'apprenticeship', pupils: 9, percentage: 9.4, status: 'published' },
{ category: 'employment', pupils: 13, percentage: 13.5, status: 'published' },
{ category: 'not_sustained', pupils: 6, percentage: 6.3, status: 'published' },
],
},
},
};
describe('Post16DestinationsSection', () => {
it('names the Year 13 cohort, not Year 11', () => {
const { container } = render(<Post16DestinationsSection destinations={phase} />);
expect(container.textContent).toMatch(/Year 13/);
expect(container.textContent).not.toMatch(/Year 11/);
});
it('reports higher education destinations', () => {
render(<Post16DestinationsSection destinations={phase} />);
expect(screen.getByText(/UK higher education/i)).toBeInTheDocument();
});
it('uses its own anchor so the nav does not collide with After Year 11', () => {
const { container } = render(<Post16DestinationsSection destinations={phase} />);
expect(container.querySelector('#post16-destinations')).toBeTruthy();
expect(container.querySelector('#destinations')).toBeNull();
});
it('renders nothing when no group carries categories', () => {
const empty: DestinationPhase = { cohort_year: '2022/23', groups: {} };
const { container } = render(<Post16DestinationsSection destinations={empty} />);
expect(container.firstChild).toBeNull();
});
});
@@ -0,0 +1,38 @@
/**
* The trail has to be written by something, and it has to be written on every
* route — not only the ones that happen to track an event.
*/
import { render } from '@testing-library/react';
const recordVisitedPath = jest.fn();
let pathname = '/schools/brentwood';
jest.mock('next/navigation', () => ({ usePathname: () => pathname }));
jest.mock('@/lib/analytics', () => ({
recordVisitedPath: (p: string) => recordVisitedPath(p),
}));
// eslint-disable-next-line @typescript-eslint/no-var-requires
const { RouteTrail } = require('@/components/RouteTrail');
describe('RouteTrail', () => {
beforeEach(() => recordVisitedPath.mockClear());
it('records the page it is mounted on', () => {
render(<RouteTrail />);
expect(recordVisitedPath).toHaveBeenCalledWith('/schools/brentwood');
});
it('records each new route as the user moves through the app', () => {
const { rerender } = render(<RouteTrail />);
pathname = '/school/115429-brentwood-school';
rerender(<RouteTrail />);
expect(recordVisitedPath).toHaveBeenLastCalledWith(
'/school/115429-brentwood-school');
});
it('renders nothing, so it can sit anywhere in the layout', () => {
const { container } = render(<RouteTrail />);
expect(container).toBeEmptyDOMElement();
});
});
@@ -0,0 +1,50 @@
import { render, screen } from '@testing-library/react';
import { SuggestList, suggestOptionId } from '@/components/SuggestList';
const ROWS = [
{ urn: 1, school_name: "St Mary's Primary", local_authority: 'Camden',
postcode: 'NW1 1AA', phase: 'Primary', school_type: 'Voluntary aided school' },
{ urn: 2, school_name: "St Mary's Primary", local_authority: 'Barnet',
postcode: 'EN5 2AA', phase: 'Primary', school_type: 'Community school' },
];
describe('SuggestList', () => {
it('is a listbox of options', () => {
render(<SuggestList id="s" suggestions={ROWS} activeIndex={-1}
onPick={() => {}} onHover={() => {}} />);
expect(screen.getByRole('listbox')).toBeInTheDocument();
expect(screen.getAllByRole('option')).toHaveLength(2);
});
it('shows the local authority, which is what tells two schools apart', () => {
// Both rows are "St Mary's Primary". Without the authority the list is
// unusable for exactly the query autosuggest exists to serve.
render(<SuggestList id="s" suggestions={ROWS} activeIndex={-1}
onPick={() => {}} onHover={() => {}} />);
expect(screen.getByText('Camden')).toBeInTheDocument();
expect(screen.getByText('Barnet')).toBeInTheDocument();
});
it('marks only the active option selected', () => {
render(<SuggestList id="s" suggestions={ROWS} activeIndex={1}
onPick={() => {}} onHover={() => {}} />);
const options = screen.getAllByRole('option');
expect(options[0]).toHaveAttribute('aria-selected', 'false');
expect(options[1]).toHaveAttribute('aria-selected', 'true');
});
it('gives each option the id the input will point at', () => {
// aria-activedescendant on the input has to name a real element id, or
// a screen reader announces nothing as the user arrows through.
render(<SuggestList id="s" suggestions={ROWS} activeIndex={0}
onPick={() => {}} onHover={() => {}} />);
expect(screen.getAllByRole('option')[0]).toHaveAttribute(
'id', suggestOptionId('s', 0));
});
it('renders nothing when there is nothing to suggest', () => {
const { container } = render(<SuggestList id="s" suggestions={[]}
activeIndex={-1} onPick={() => {}} onHover={() => {}} />);
expect(container).toBeEmptyDOMElement();
});
});
@@ -0,0 +1,45 @@
import { render } from '@testing-library/react';
import { TrackPlaceView } from '@/components/places/TrackPlaceView';
const trackMock = jest.fn();
jest.mock('@/lib/analytics', () => ({
track: (...args: unknown[]) => trackMock(...args),
getNavigationSource: () => 'search',
}));
describe('TrackPlaceView', () => {
beforeEach(() => trackMock.mockClear());
it('reports which kind of location page was viewed', () => {
/*
* `kind` is the reason this event exists. Whether to keep investing in the
* location layer turns on which *sort* of page earns engagement — towns,
* authorities or postcode districts — and a bare pageview cannot say,
* because all four families share the /schools/ prefix.
*/
render(<TrackPlaceView kind="authority" slug="kent" count={412} />);
expect(trackMock).toHaveBeenCalledWith('place_viewed', {
kind: 'authority', slug: 'kent', phase: 'all',
school_count: 412, from: 'search',
});
});
it('names the phase when the page is a phase variant', () => {
render(<TrackPlaceView kind="town" slug="brentwood" count={29} phase="primary" />);
expect(trackMock).toHaveBeenCalledWith('place_viewed',
expect.objectContaining({ phase: 'primary' }));
});
it('fires once, not once per render', () => {
const { rerender } = render(
<TrackPlaceView kind="town" slug="brentwood" count={29} />);
rerender(<TrackPlaceView kind="town" slug="brentwood" count={29} />);
expect(trackMock).toHaveBeenCalledTimes(1);
});
it('renders nothing', () => {
const { container } = render(
<TrackPlaceView kind="town" slug="brentwood" count={29} />);
expect(container).toBeEmptyDOMElement();
});
});
@@ -0,0 +1,192 @@
import fs from 'fs';
import path from 'path';
/**
* Guards against light-theme-only CSS.
*
* The site themes entirely through tokens redefined under
* `@media (prefers-color-scheme: dark)`. A hardcoded colour therefore does not
* fail loudly — it renders perfectly in the theme it was written for and
* quietly wrongly in the other, which nobody sees unless they happen to be in
* dark mode when they look.
*
* Both rules below are drawn from real defects in SchoolHeroMap.module.css,
* found by eye rather than by any test:
*
* - the map's fade to the header ramped through hardcoded white and landed on
* `var(--bg-card)`. Invisible in light; a bright band across the full width
* of a near-black card in dark.
* - the controls floating over the map paired a hardcoded white background
* with `color: var(--text-primary)`, which resolves to #E9EEF0 in dark —
* near-white text on a near-white button.
*/
const COMPONENTS = path.join(__dirname, '..', '..', 'components');
function stylesheets(dir: string): string[] {
return fs.readdirSync(dir, { withFileTypes: true }).flatMap((entry) => {
const full = path.join(dir, entry.name);
if (entry.isDirectory()) return stylesheets(full);
return entry.name.endsWith('.module.css') ? [full] : [];
});
}
/** Innermost `selector { body }` pairs. Nested at-rules never match as rules,
* because their body contains braces.
*
* Comments are stripped before matching rather than after, so that the whole
* selector survives. Taking only its last line — which is what stripping a
* leading comment used to require — silently discarded every selector in a
* grouped rule but the final one, and a safety guard that cannot see half its
* input fails open. */
function rules(css: string): Array<{ selector: string; body: string }> {
const bare = css.replace(/\/\*[\s\S]*?\*\//g, '');
return Array.from(bare.matchAll(/([^{}]+)\{([^{}]*)\}/g), (m) => ({
selector: m[1].trim().replace(/\s*\n\s*/g, ' '),
body: m[2],
}));
}
const HARDCODED_WHITE_BG = /background[^;]*(?:255,\s*255,\s*255|#fff\b|#ffffff\b)/i;
const THEMED_COLOR = /(?:^|[^-])color:\s*var\(--/;
const files = stylesheets(COMPONENTS);
/** Component sources, for the third-party-surface rule below. */
function sources(dir: string): string[] {
return fs.readdirSync(dir, { withFileTypes: true }).flatMap((entry) => {
const full = path.join(dir, entry.name);
if (entry.isDirectory()) return sources(full);
return entry.name.endsWith('.tsx') ? [full] : [];
});
}
describe('dark-theme safety', () => {
it('finds stylesheets to check', () => {
expect(files.length).toBeGreaterThan(0);
});
it('never pairs a hardcoded white background with a themed text colour', () => {
const offenders = files.flatMap((file) =>
rules(fs.readFileSync(file, 'utf8'))
.filter((r) => HARDCODED_WHITE_BG.test(r.body) && THEMED_COLOR.test(r.body))
.map((r) => `${path.relative(COMPONENTS, file)} ${r.selector}`));
// Either the surface follows the theme and so should the text, or it does
// not and the text must be literal too. Mixing them is how near-white text
// ends up on a near-white button.
expect(offenders).toEqual([]);
});
it('never fades to a themed colour through a hardcoded one', () => {
const offenders = files.flatMap((file) =>
rules(fs.readFileSync(file, 'utf8'))
.filter((r) => /linear-gradient/.test(r.body)
&& /var\(--bg-(card|primary|secondary)\)/.test(r.body)
&& /255,\s*255,\s*255|#fff\b/i.test(r.body))
.map((r) => `${path.relative(COMPONENTS, file)} ${r.selector}`));
// A gradient that lands on a token has to be made of that token, or the
// ramp and its destination disagree in one theme. Use the matching
// `--*-rgb` token for the transparent stops.
expect(offenders).toEqual([]);
});
});
/**
* The same defect one stylesheet further out.
*
* The rules above scan our own CSS modules. They cannot see a surface painted
* by a third-party sheet: leaflet.css hardcodes `background: white` on
* `.leaflet-popup-content-wrapper` and `.leaflet-popup-tip`, and
* LeafletMapInner builds its popup as an HTML string with inline
* `color: var(--text-primary)`. Neither half lives in a .module.css, so the
* module scan passed while dark mode rendered #E9EEF0 on #FFFFFF — 1.17:1,
* with the school name and the headline figure effectively invisible.
*
* globals.css already pulls the rest of Leaflet's chrome onto the tokens (the
* attribution bar, the zoom controls) for exactly this reason. The popup was
* simply missed.
*/
describe('third-party surfaces under themed text', () => {
const GLOBALS = path.join(__dirname, '..', '..', 'app', '(frontend)', 'globals.css');
/** Leaflet surfaces our own code writes token-coloured text onto. */
const LEAFLET_POPUP_SURFACES = [
'.leaflet-popup-content-wrapper',
'.leaflet-popup-tip',
];
it('still finds a component painting themed text into a Leaflet popup', () => {
// Guards the rule below against passing vacuously if the popups are ever
// rewritten as React components rather than HTML strings.
const themed = sources(COMPONENTS).filter((file) => {
const src = fs.readFileSync(file, 'utf8');
return /bindPopup\(/.test(src) && /color:var\(--|color: var\(--/.test(src);
});
expect(themed.length).toBeGreaterThan(0);
});
it('themes the Leaflet popup surface, because the text on it is themed', () => {
const globals = rules(fs.readFileSync(GLOBALS, 'utf8'));
const unthemed = LEAFLET_POPUP_SURFACES.filter((surface) => {
const rule = globals.find((r) => r.selector.includes(surface));
return !rule || !/background[^;]*var\(--/.test(rule.body);
});
// Leaflet's white is not a colour this site owns. Either the surface
// follows the theme or the text on it must be literal — and the text is
// already themed.
expect(unthemed).toEqual([]);
});
it('never puts a literal white label on a themed fill', () => {
/*
* The mirror image of the module-CSS rule above, and the half of the popup
* that theming the card does not reach. "View Details" is
* `background:var(--status-above);color:white`; --status-above is #36743F
* in light but #7FCB8A in dark, so the label went from 5.63:1 to 1.94:1.
*
* --text-inverse is the token for ink on a saturated fill — #FFFFFF in
* light, #111A20 in dark — and the popup's Ofsted badge already uses it.
*/
const offenders = sources(COMPONENTS).flatMap((file) => {
const src = fs.readFileSync(file, 'utf8');
return Array.from(
src.matchAll(/background:\s*var\(--[^;"']*;[^"']*?color:\s*(white|#fff\b|#ffffff\b)/gi),
() => path.relative(COMPONENTS, file));
});
expect(offenders).toEqual([]);
});
});
/**
* Destination measures add the first new colour family since the palette was
* set. The tokens have to exist in both blocks or the section renders one
* theme's fills on the other theme's ground — the exact failure the suite
* above exists to catch, but for tokens rather than literals.
*/
describe('destination tokens', () => {
const css = fs.readFileSync(
path.join(__dirname, '..', '..', 'app', '(frontend)', 'globals.css'), 'utf8');
const TOKENS = [
'--dest-sixthform', '--dest-sfcollege', '--dest-fecollege',
'--dest-apprentice', '--dest-employment', '--dest-none', '--dest-none-hatch',
];
const DARK_AT = css.indexOf('@media (prefers-color-scheme: dark)');
it('defines every destination token in the light palette', () => {
const light = css.slice(0, DARK_AT);
expect(TOKENS.filter((t) => !light.includes(`${t}:`))).toEqual([]);
});
it('redefines every destination token for dark', () => {
const dark = css.slice(DARK_AT);
expect(TOKENS.filter((t) => !dark.includes(`${t}:`))).toEqual([]);
});
});
@@ -0,0 +1,108 @@
import fs from 'fs';
import path from 'path';
/**
* The hero search and the results filter bar are the same component in two
* costumes. `.filterBar` is the card — background, border, shadow, padding —
* and `.heroMode` strips all of it so the search sits directly on the hero
* panel.
*
* Both selectors have specificity (0,1,0), so **source order decides**, and
* `.heroMode` only wins because it is declared immediately after. Any later
* bare `.filterBar` rule — which in practice means one inside a media query —
* silently wins instead, and the hero grows a card's padding back.
*
* That is exactly what happened: `@media (max-width: 768px) { .filterBar {
* padding: 0.875rem } }` re-added 14px in hero mode, indenting the search box,
* the hint and the location link 14px past the headline above them and costing
* the search field 28px of width on a 390px screen. The two rules directly
* below it in the same block were correctly written as
* `.filterBar:not(.heroMode)`; this one was missed, and nothing caught it
* because the result is a plausible-looking layout rather than a broken one.
*/
const CSS = path.join(__dirname, '..', '..', 'components', 'FilterBar.module.css');
/** Properties `.heroMode` resets. A later bare `.filterBar` rule setting any
* of these puts the card back on the hero. */
const RESET_BY_HERO_MODE = [
'background', 'border', 'border-radius', 'box-shadow', 'padding',
];
/**
* Comments are stripped before anything is parsed.
*
* A `{` or `}` inside a comment would otherwise desynchronise the brace walk
* below and the rule regex alike, and the selector text captured for each rule
* would carry the preceding comment along with it.
*/
function withoutComments(css: string): string {
return css.replace(/\/\*[\s\S]*?\*\//g, '');
}
/**
* The individual selectors in a rule's prelude.
*
* Split on commas, because a selector list is a list: `.filterBar, .other { }`
* applies to `.filterBar` just as surely as `.filterBar { }` does, and an
* earlier version of this guard compared the whole prelude against the literal
* string '.filterBar' — so writing the regression as a comma list, or across
* two lines, would have walked straight past it.
*/
function selectorsOf(prelude: string): string[] {
return prelude.split(',').map((sel) => sel.trim().replace(/\s+/g, ' '))
.filter(Boolean);
}
function mediaQueryBodies(css: string): string[] {
const bodies: string[] = [];
const re = /@media[^{]*\{/g;
let m: RegExpExecArray | null;
while ((m = re.exec(css)) !== null) {
// Walk braces from the opening one to find this at-rule's whole body.
let depth = 1;
let i = m.index + m[0].length;
const start = i;
while (i < css.length && depth > 0) {
if (css[i] === '{') depth++;
else if (css[i] === '}') depth--;
i++;
}
bodies.push(css.slice(start, i - 1));
}
return bodies;
}
describe('FilterBar hero-mode scoping', () => {
const css = withoutComments(fs.readFileSync(CSS, 'utf8'));
it('confirms heroMode still resets the card, which is what makes this matter', () => {
const hero = css.match(/\.heroMode\s*\{([^}]*)\}/);
expect(hero).not.toBeNull();
expect(hero![1]).toMatch(/padding:\s*0/);
});
it('never re-applies card styling to the hero from inside a media query', () => {
const offenders: string[] = [];
for (const body of mediaQueryBodies(css)) {
for (const rule of body.matchAll(/([^{}]+)\{([^{}]*)\}/g)) {
// Only a *bare* .filterBar is dangerous, and it is dangerous wherever
// it appears in a selector list. Scoped variants
// (`.filterBar:not(.heroMode)`) and descendants are fine.
const selectors = selectorsOf(rule[1]);
if (!selectors.includes('.filterBar')) continue;
for (const prop of RESET_BY_HERO_MODE) {
if (new RegExp(`(^|[;\\s])${prop}\\s*:`).test(rule[2])) {
offenders.push(`${rule[1].trim()} sets ${prop}`);
}
}
}
}
// Fix by scoping the rule as `.filterBar:not(.heroMode)`, the way the
// neighbouring rules in the same block already are.
expect(offenders).toEqual([]);
});
});
@@ -98,6 +98,17 @@ describe('secondary detail page', () => {
expect(screen.getByText(/has not published a cut-off distance/)).toBeInTheDocument(); expect(screen.getByText(/has not published a cut-off distance/)).toBeInTheDocument();
}); });
it('makes no claim about publication when the feature is switched off', () => {
// Absent, not null. The API omits the key entirely while the
// admission_distance flag is off, and "Islington has not published a
// cut-off distance" is then a statement about us, not about Islington —
// false wherever the authority does publish one.
renderSecondarySchoolDetail({ ...secondaryFixture, admissionDistance: undefined });
expect(screen.queryByText(/has not published a cut-off distance/)).not.toBeInTheDocument();
expect(screen.queryByText(/Contact the admissions authority/)).not.toBeInTheDocument();
});
}); });
// ── The Distance section ─────────────────────────────────────────────── // ── The Distance section ───────────────────────────────────────────────
@@ -0,0 +1,73 @@
import { renderHook, act, waitFor } from '@testing-library/react';
import { useSchoolSuggest } from '@/hooks/useSchoolSuggest';
const realFetch = global.fetch;
function mockFetch(rows: unknown[], delayMs = 0) {
global.fetch = jest.fn(async (_url: unknown, init?: { signal?: AbortSignal }) => {
if (delayMs) {
await new Promise((resolve, reject) => {
const t = setTimeout(resolve, delayMs);
init?.signal?.addEventListener('abort', () => {
clearTimeout(t);
reject(Object.assign(new Error('aborted'), { name: 'AbortError' }));
});
});
}
return { ok: true, json: async () => ({ suggestions: rows }) };
}) as unknown as typeof fetch;
}
const ROW = {
urn: 1, school_name: 'Brecknock Primary School', local_authority: 'Camden',
postcode: 'NW1 1AA', phase: 'Primary', school_type: 'Community school',
};
describe('useSchoolSuggest', () => {
beforeEach(() => { jest.useFakeTimers(); });
afterEach(() => { jest.useRealTimers(); global.fetch = realFetch; });
it('does not fetch below the minimum query length', () => {
mockFetch([ROW]);
renderHook(() => useSchoolSuggest('b', true));
act(() => { jest.advanceTimersByTime(500); });
expect(global.fetch).not.toHaveBeenCalled();
});
it('does not fetch at all when disabled', () => {
// The flag being off must mean no request, not a hidden dropdown.
mockFetch([ROW]);
renderHook(() => useSchoolSuggest('brecknock', false));
act(() => { jest.advanceTimersByTime(500); });
expect(global.fetch).not.toHaveBeenCalled();
});
it('debounces rather than firing per keystroke', () => {
mockFetch([ROW]);
const { rerender } = renderHook(
({ q }) => useSchoolSuggest(q, true), { initialProps: { q: 'br' } });
rerender({ q: 'bre' });
rerender({ q: 'brec' });
act(() => { jest.advanceTimersByTime(199); });
expect(global.fetch).not.toHaveBeenCalled();
act(() => { jest.advanceTimersByTime(2); });
expect(global.fetch).toHaveBeenCalledTimes(1);
});
it('opens with results once they arrive', async () => {
mockFetch([ROW]);
const { result } = renderHook(() => useSchoolSuggest('brecknock', true));
act(() => { jest.advanceTimersByTime(200); });
await waitFor(() => expect(result.current.suggestions).toHaveLength(1));
expect(result.current.open).toBe(true);
});
it('close() hides the list without clearing the query', async () => {
mockFetch([ROW]);
const { result } = renderHook(() => useSchoolSuggest('brecknock', true));
act(() => { jest.advanceTimersByTime(200); });
await waitFor(() => expect(result.current.open).toBe(true));
act(() => { result.current.close(); });
expect(result.current.open).toBe(false);
});
});
+158
View File
@@ -0,0 +1,158 @@
import { getNavigationSource } from '@/lib/analytics';
/** jsdom's document.referrer is read-only; redefining it is the way in. */
function referrer(url: string) {
Object.defineProperty(document, 'referrer', { value: url, configurable: true });
}
const ORIGIN = 'http://localhost';
describe('getNavigationSource', () => {
afterEach(() => referrer(''));
it('attributes a visit from a location page to the place layer', () => {
/*
* The one this was added for.
*
* W2 published ~3,900 location pages whose entire purpose is to funnel
* search traffic onto school pages. Before this case existed they fell
* through to 'direct' — so the location layer's contribution was not
* merely missing from the funnel, it was being counted in the bucket you
* read as "typed the URL". The measurement that decides whether W2 worked
* was confidently reporting the wrong answer.
*/
referrer(`${ORIGIN}/schools/barnet`);
expect(getNavigationSource()).toBe('place');
});
it.each([
['/schools/authority/kent', 'authority'],
['/schools/near/sw11', 'outcode'],
['/schools/brentwood/primary', 'phase variant'],
])('covers %s (%s)', (path) => {
referrer(`${ORIGIN}${path}`);
expect(getNavigationSource()).toBe('place');
});
it('still calls a school page "detail", one character away', () => {
// /school/ and /schools/ differ by one letter and mean different things.
// A prefix test written in the wrong order silently merges them.
referrer(`${ORIGIN}/school/100010-brecknock-primary-school`);
expect(getNavigationSource()).toBe('detail');
});
it.each([
['/', 'search'],
['/rankings', 'rankings'],
['/compare?urns=1,2', 'compare'],
])('leaves %s attributed as %s', (path, expected) => {
referrer(`${ORIGIN}${path}`);
expect(getNavigationSource()).toBe(expected);
});
it('treats an external referrer as direct', () => {
// Umami records the real referrer on the pageview; this field is only
// about internal navigation.
referrer('https://www.google.com/search?q=schools+in+barnet');
expect(getNavigationSource()).toBe('direct');
});
it('treats no referrer as direct', () => {
referrer('');
expect(getNavigationSource()).toBe('direct');
});
});
/*
* The defect the existing suite could not see.
*
* Every test above sets document.referrer, which the browser writes only when
* a *document* loads. Every internal navigation in this app is an App Router
* soft navigation — history.pushState, no new document — so document.referrer
* keeps naming whatever opened the tab for the whole session. Verified on
* staging: /schools/brentwood → click a school → URL changes to /school/…
* and document.referrer is still "".
*
* So `from` reported 'direct' for essentially every in-app journey, and the
* suite passed because it only ever exercised the full-page-load path.
*/
function freshAnalytics() {
let mod!: typeof import('@/lib/analytics');
jest.isolateModules(() => {
mod = require('@/lib/analytics');
});
return mod;
}
function at(path: string) {
window.history.pushState({}, '', path);
}
describe('getNavigationSource across a soft navigation', () => {
afterEach(() => {
referrer('');
at('/');
});
it('attributes a school view to the place page the user actually came from', () => {
const { recordVisitedPath, getNavigationSource: source } = freshAnalytics();
at('/schools/brentwood');
recordVisitedPath('/schools/brentwood');
at('/school/115429-brentwood-school');
recordVisitedPath('/school/115429-brentwood-school');
expect(source()).toBe('place');
});
it('does not depend on whether the new path was recorded first', () => {
// The trail is written by a layout-level effect and read by a page-level
// one. React orders those by tree position, which is not a contract worth
// resting a measurement on, so the answer must be the same either way.
const { recordVisitedPath, getNavigationSource: source } = freshAnalytics();
recordVisitedPath('/rankings');
at('/school/115429-brentwood-school');
expect(source()).toBe('rankings');
});
it('names the previous page, not the current one, when both are schools', () => {
const { recordVisitedPath, getNavigationSource: source } = freshAnalytics();
at('/school/100010-brecknock-primary-school');
recordVisitedPath('/school/100010-brecknock-primary-school');
at('/school/115429-brentwood-school');
recordVisitedPath('/school/115429-brentwood-school');
expect(source()).toBe('detail');
});
it('looks past a return visit to the page the user came back from', () => {
const { recordVisitedPath, getNavigationSource: source } = freshAnalytics();
for (const p of ['/schools/brentwood', '/school/115429-brentwood-school',
'/schools/brentwood']) {
at(p);
recordVisitedPath(p);
}
expect(source()).toBe('detail');
});
it('falls back to the referrer on a real document load, where it is true', () => {
// A fresh module is a fresh document: nothing has been recorded, and
// document.referrer is meaningful again.
const { getNavigationSource: source } = freshAnalytics();
at('/school/115429-brentwood-school');
referrer(`${ORIGIN}/schools/barnet`);
expect(source()).toBe('place');
});
it('still reads an arrival from outside as direct', () => {
const { recordVisitedPath, getNavigationSource: source } = freshAnalytics();
at('/schools/brentwood');
recordVisitedPath('/schools/brentwood');
referrer('https://www.google.com/search?q=schools+in+brentwood');
expect(source()).toBe('direct');
});
});
@@ -0,0 +1,87 @@
import {
canAggregate, aggregateCells,
canRenderBar, toBarSegments, CARD_GROUPS,
type DestinationCell, type DestinationGroup, type DestinationCategory,
} from '@/lib/destinations';
const pub = (category: DestinationCategory, pupils: number, cohort: number): DestinationCell => ({
category, pupils, percentage: (pupils / cohort) * 100, status: 'published',
});
const sup = (category: DestinationCategory): DestinationCell => ({
category, pupils: null, percentage: null, status: 'suppressed',
});
const fullGroup = (): DestinationGroup => ({
cohort: 180,
cells: [
pub('school_sixth_form', 75, 180), pub('sixth_form_college', 21, 180),
pub('further_education', 55, 180), pub('other_education', 6, 180),
pub('apprenticeship', 8, 180), pub('employment', 6, 180),
pub('not_sustained', 5, 180), pub('not_captured', 4, 180),
],
});
describe('canAggregate — R2, computing from components', () => {
it('allows a sum when every component is published', () => {
expect(canAggregate([pub('apprenticeship', 8, 180), pub('employment', 6, 180)])).toBe(true);
});
it('refuses a sum when any component is suppressed', () => {
expect(canAggregate([pub('apprenticeship', 8, 180), sup('employment')])).toBe(false);
});
it('refuses a sum when every component is suppressed', () => {
expect(canAggregate([sup('apprenticeship'), sup('employment')])).toBe(false);
});
});
describe('aggregateCells', () => {
it('sums published cells and derives a percentage from the cohort', () => {
expect(aggregateCells([pub('apprenticeship', 8, 180), pub('employment', 6, 180)], 180))
.toEqual({ pupils: 14, percentage: (14 / 180) * 100 });
});
it('returns null rather than a partial sum when a component is suppressed', () => {
expect(aggregateCells([pub('apprenticeship', 8, 180), sup('employment')], 180)).toBeNull();
});
});
describe('canRenderBar — R1', () => {
it('allows a bar when the whole group is published', () => {
expect(canRenderBar(fullGroup())).toBe(true);
});
it('refuses a bar when a single category is suppressed', () => {
const g = fullGroup();
g.cells[1] = sup('sixth_form_college');
expect(canRenderBar(g)).toBe(false);
});
});
describe('toBarSegments', () => {
it('derives widths from counts, not from rounded percentages', () => {
const segs = toBarSegments(fullGroup());
expect(segs).toHaveLength(8);
expect(segs[0].widthPct).toBeCloseTo((75 / 180) * 100, 10);
expect(segs.reduce((a, s) => a + s.widthPct, 0)).toBeCloseTo(100, 6);
});
it('throws rather than silently leaving a gap when the group is suppressed', () => {
const g = fullGroup();
g.cells[1] = sup('sixth_form_college');
expect(() => toBarSegments(g)).toThrow(/suppressed/i);
});
});
describe('CARD_GROUPS', () => {
it('partitions every destination category exactly once, plus the absence', () => {
const grouped = Object.values(CARD_GROUPS).flat();
expect(new Set(grouped).size).toBe(grouped.length);
expect(grouped).toEqual(expect.arrayContaining([
'school_sixth_form', 'sixth_form_college', 'further_education',
'other_education', 'apprenticeship', 'employment',
]));
expect(grouped).not.toContain('not_sustained');
expect(grouped).not.toContain('not_captured');
});
});
+58
View File
@@ -0,0 +1,58 @@
import { getFlags, FLAGS_REVALIDATE } from '@/lib/flags';
// jsdom provides no global fetch, so there is nothing for jest.spyOn to attach
// to — assign it and restore the original afterwards. This is the first test
// here to mock fetch; later ones should follow this shape.
const realFetch = global.fetch;
function mockFetch(impl: () => Promise<unknown>) {
global.fetch = jest.fn(impl) as unknown as typeof fetch;
}
describe('getFlags', () => {
afterEach(() => { global.fetch = realFetch; });
it('returns the flags the API reports', async () => {
mockFetch(async () => ({
ok: true,
json: async () => ({ admission_distance: true }),
}));
await expect(getFlags()).resolves.toEqual({ admission_distance: true });
});
it('returns no flags rather than throwing when the API is down', async () => {
// A page that cannot read flags must render everything dark, not 500.
// Fail-closed is the same direction as the backend's default.
mockFetch(async () => { throw new Error('ECONNREFUSED'); });
await expect(getFlags()).resolves.toEqual({});
});
it('returns no flags rather than throwing on a non-200', async () => {
mockFetch(async () => ({ ok: false, status: 503 }));
await expect(getFlags()).resolves.toEqual({});
});
/*
* Reading a flag pins the calling route's ISR floor: Next uses the LOWEST
* revalidate among a route's fetches for the whole route. That is why the
* revalidate is an argument rather than the constant.
*
* Every SEO route here declares `revalidate = 604800`. A gate that read
* flags at the 300s default would drop the whole school and place corpus
* from a weekly cache to a 5-minute one, which is a large origin-load
* regression to pay for a feature flag.
*/
it('reads at the 300s floor by default', async () => {
mockFetch(async () => ({ ok: true, json: async () => ({}) }));
await getFlags();
expect((global.fetch as jest.Mock).mock.calls[0][1])
.toEqual({ next: { revalidate: FLAGS_REVALIDATE } });
});
it('lets a caller pass its own route floor instead', async () => {
mockFetch(async () => ({ ok: true, json: async () => ({}) }));
await getFlags(604800);
expect((global.fetch as jest.Mock).mock.calls[0][1])
.toEqual({ next: { revalidate: 604800 } });
});
});
@@ -0,0 +1,83 @@
import { computeSecondaryFlags, buildSecondaryNavItems } from '@/lib/schoolSections';
import type { School, SchoolDestinations } from '@/lib/types';
const schoolInfo = {
urn: 137083, school_name: 'Northbrook Academy', phase: 'Secondary',
has_sixth_form: true,
} as unknown as School;
const base = { schoolInfo, yearlyData: [], deprivation: null, finance: null };
const phase = (categories = 1) => ({
cohort_year: '2022/23',
groups: {
all: {
cohort: 180,
categories: Array.from({ length: categories }, () => ({
category: 'school_sixth_form' as const,
pupils: 75, percentage: 41.7, status: 'published' as const,
})),
},
},
});
const ks4Only: SchoolDestinations = { ks4: phase(), ks5: null };
const both: SchoolDestinations = { ks4: phase(), ks5: phase() };
describe('computeSecondaryFlags — destinations', () => {
it('flags KS4 destinations when the block carries categories', () => {
const flags = computeSecondaryFlags({ ...base, destinations: ks4Only });
expect(flags.hasKs4Destinations).toBe(true);
expect(flags.hasKs5Destinations).toBe(false);
});
it('flags both phases when both are present', () => {
const flags = computeSecondaryFlags({ ...base, destinations: both });
expect(flags.hasKs4Destinations).toBe(true);
expect(flags.hasKs5Destinations).toBe(true);
});
it('flags neither when the block is absent', () => {
const flags = computeSecondaryFlags({ ...base, destinations: null });
expect(flags.hasKs4Destinations).toBe(false);
expect(flags.hasKs5Destinations).toBe(false);
});
it('does not flag a phase whose groups carry no categories', () => {
const empty: SchoolDestinations = {
ks4: { cohort_year: '2022/23', groups: {} }, ks5: null,
};
expect(computeSecondaryFlags({ ...base, destinations: empty }).hasKs4Destinations)
.toBe(false);
});
it('does not flag a phase whose only group has an empty category list', () => {
const empty: SchoolDestinations = { ks4: phase(0), ks5: null };
expect(computeSecondaryFlags({ ...base, destinations: empty }).hasKs4Destinations)
.toBe(false);
});
});
describe('buildSecondaryNavItems — destinations', () => {
const navInput = {
ofsted: null, admissions: null, admissionDistance: null,
hasLocation: false, yearlyDataLength: 0,
};
it('adds both entries, after GCSEs', () => {
const flags = computeSecondaryFlags({ ...base, destinations: both });
const ids = buildSecondaryNavItems({ ...flags, hasResults: true }, navInput)
.map(i => i.id);
expect(ids).toContain('destinations');
expect(ids).toContain('post16-destinations');
expect(ids.indexOf('destinations')).toBeGreaterThan(ids.indexOf('gcse'));
expect(ids.indexOf('post16-destinations')).toBe(ids.indexOf('destinations') + 1);
});
it('adds no entry for a phase that will not render — the nav must not link to a missing anchor', () => {
const flags = computeSecondaryFlags({ ...base, destinations: null });
const ids = buildSecondaryNavItems(flags, navInput).map(i => i.id);
expect(ids).not.toContain('destinations');
expect(ids).not.toContain('post16-destinations');
});
});
+26
View File
@@ -13,6 +13,8 @@ import {
metricKind, metricKind,
shortName, shortName,
computeYBounds, computeYBounds,
formatAgeRange,
formatAgeSpan,
} from '@/lib/utils'; } from '@/lib/utils';
describe('formatPercentage', () => { describe('formatPercentage', () => {
@@ -320,3 +322,27 @@ describe('shortName', () => {
expect(shortName('A'.repeat(30), 10)).toBe('AAAAAAAAA…'); expect(shortName('A'.repeat(30), 10)).toBe('AAAAAAAAA…');
}); });
}); });
describe('formatAgeSpan', () => {
it('normalises a hyphenated range to an en dash, without a label', () => {
// The place table carries "Ages" in the column heading, so repeating it
// in every cell is noise. formatAgeRange keeps the label for the contexts
// that have no heading to hang it on.
expect(formatAgeSpan('4-11')).toBe('4–11');
});
it('leaves a range it does not recognise alone rather than mangling it', () => {
expect(formatAgeSpan('3-19 (SEN)')).toBe('3-19 (SEN)');
});
it('returns an empty string for a missing range', () => {
expect(formatAgeSpan(null)).toBe('');
expect(formatAgeSpan(undefined)).toBe('');
});
});
describe('formatAgeRange', () => {
it('keeps its label, so the two helpers stay distinguishable', () => {
expect(formatAgeRange('4-11')).toBe('Ages 4–11');
});
});
@@ -0,0 +1,70 @@
/**
* Payload is ESM-only and next/jest will not transform it, so the collections
* cannot be imported and their sanitised config inspected here (see
* lib/payloadRoutes.ts for the full reasoning). These assert the source of the
* collection definitions instead — enough to catch the settings whose loss is
* silent, and cheap. Behaviour is proved by the e2e journeys against staging.
*/
import fs from 'fs';
import path from 'path';
const read = (file: string) =>
fs.readFileSync(path.join(__dirname, '..', '..', 'collections', file), 'utf8');
const POSTS = read('Posts.ts');
const MEDIA = read('Media.ts');
const CONFIG = fs.readFileSync(
path.join(__dirname, '..', '..', 'payload.config.ts'),
'utf8',
);
describe('posts collection', () => {
it('supports drafts, so saving is not publishing', () => {
expect(POSTS).toMatch(/drafts:\s*true/);
});
it('has a unique, indexed slug for stable URLs', () => {
const slugField = POSTS.slice(POSTS.indexOf("name: 'slug'"));
expect(slugField).toMatch(/unique:\s*true/);
expect(slugField).toMatch(/index:\s*true/);
});
it('hides drafts from anonymous readers at the access layer', () => {
// Payload's docs are explicit: "The `draft` argument alone does not
// restrict documents with _status: 'draft' from being returned by the
// API." The blog pages' where-clause is not enforcement — a direct GET
// /cms-api/posts would return unpublished drafts to anyone. Access
// control returning a query constraint is the only thing that stops it.
expect(POSTS).toMatch(/_status:\s*\{\s*equals:\s*'published'\s*\}/);
expect(POSTS).toMatch(/if\s*\(req\.user\)\s*return true/);
});
it('revalidates the post page when a post changes or is deleted', () => {
// /blog/[slug] is ISR — generated on first request and cached — so an edit
// to an already-published post would otherwise not appear until the
// revalidate window expired, up to an hour of a writer concluding that
// saving is broken. The index and feeds are force-dynamic and need no hook.
expect(POSTS).toContain('afterChange');
expect(POSTS).toContain('afterDelete');
expect(POSTS).toMatch(/revalidatePath\(`\/blog\/\$\{[^}]+\}`\)/);
});
});
describe('media collection', () => {
it('writes uploads to the mounted volume, by absolute path', () => {
// Must match the payload_media mount in docker-compose.portainer.yml.
// Payload 3 requires staticDir to be absolute.
expect(MEDIA).toMatch(/staticDir:\s*'\/app\/media'/);
});
it('requires alt text on every upload', () => {
const altField = MEDIA.slice(MEDIA.indexOf("name: 'alt'"));
expect(altField).toMatch(/required:\s*true/);
});
});
describe('payload config', () => {
it('registers every collection', () => {
expect(CONFIG).toMatch(/collections:\s*\[Users,\s*Posts,\s*Media\]/);
});
});
@@ -0,0 +1,67 @@
/**
* The admin panel does not import field components directly. Payload sends the
* client a *path* for each one — a richText field's is
* `@payloadcms/richtext-lexical/rsc#RscEntryLexicalField` — and resolves it
* through this generated map. An entry that is missing from the map is not an
* error the panel reports: the field simply does not render.
*
* That failure is quietly awful, because `required: true` is enforced on the
* server regardless. A writer gets a new-post form with no Content editor and
* a save that refuses on a field they were never shown.
*
* The map is generated by `npx payload generate:importmap`, so it drifts every
* time a field or a lexical feature is added and nobody re-runs it. These
* assert the entries the current config needs.
*/
import fs from 'fs';
import path from 'path';
const MAP = fs.readFileSync(
path.join(__dirname, '..', '..', 'app', '(payload)', 'admin', 'importMap.js'),
'utf8',
);
const POSTS = fs.readFileSync(
path.join(__dirname, '..', '..', 'collections', 'Posts.ts'),
'utf8',
);
describe('admin import map', () => {
it('resolves the richText field, so Content renders in the editor', () => {
// Guarded because Posts.content is required: without this entry the field
// is invisible and the post is unsaveable.
expect(POSTS).toMatch(/type:\s*'richText'/);
expect(MAP).toContain('@payloadcms/richtext-lexical/rsc#RscEntryLexicalField');
});
it('resolves the richText cell, so the list view can render the column', () => {
expect(MAP).toContain('@payloadcms/richtext-lexical/rsc#RscEntryLexicalCell');
});
it('resolves the diff component, which the drafts UI needs', () => {
// versions.drafts is on, so the panel offers version comparison.
expect(POSTS).toMatch(/drafts:\s*true/);
expect(MAP).toContain('@payloadcms/richtext-lexical/rsc#LexicalDiffComponent');
});
it('resolves BlocksFeature, so the Callout block is insertable', () => {
expect(POSTS).toContain('BlocksFeature');
expect(MAP).toContain('@payloadcms/richtext-lexical/client#BlocksFeatureClient');
});
it('resolves the default toolbar features the editor is built with', () => {
// defaultFeatures is spread into the editor config; each one contributes a
// client component the toolbar cannot render without.
for (const feature of [
'BoldFeatureClient',
'ItalicFeatureClient',
'HeadingFeatureClient',
'LinkFeatureClient',
'UploadFeatureClient',
'UnorderedListFeatureClient',
'OrderedListFeatureClient',
'InlineToolbarFeatureClient',
]) {
expect(MAP).toContain(`@payloadcms/richtext-lexical/client#${feature}`);
}
});
});
@@ -0,0 +1,52 @@
/**
* The generated migration is schema-qualified to "payload" throughout but does
* not create that schema — `schemaName` says where tables go, it does not
* create anything. On staging and production, which have never run it, the
* whole migration fails with `schema "payload" does not exist`.
*
* The CREATE SCHEMA is therefore hand-added, which makes it exactly the kind
* of edit a regeneration silently discards. This is the guard.
*/
import fs from 'fs';
import path from 'path';
const DIR = path.join(__dirname, '..', '..', 'migrations');
function migrationFiles() {
return fs
.readdirSync(DIR)
.filter((f) => f.endsWith('.ts') && f !== 'index.ts');
}
describe('payload migrations', () => {
it('ships at least one migration, so a container has tables to find', () => {
expect(migrationFiles().length).toBeGreaterThan(0);
});
it('creates the payload schema before creating anything in it', () => {
const initial = migrationFiles().find((f) => f.includes('initial'))!;
const sql = fs.readFileSync(path.join(DIR, initial), 'utf8');
expect(sql).toMatch(/CREATE SCHEMA IF NOT EXISTS "payload"/);
// Ordering matters: the schema must be created before the first object
// that lives in it, or the migration fails on its first statement.
expect(sql.indexOf('CREATE SCHEMA IF NOT EXISTS "payload"'))
.toBeLessThan(sql.indexOf('CREATE TABLE "payload"'));
});
it('creates the tables the app queries on boot', () => {
const initial = migrationFiles().find((f) => f.includes('initial'))!;
const sql = fs.readFileSync(path.join(DIR, initial), 'utf8');
for (const table of ['users', 'posts', '_posts_v', 'media', 'payload_migrations']) {
expect(sql).toContain(`CREATE TABLE "payload"."${table}"`);
}
});
it('is wired into the adapter, so it runs on server init', () => {
const config = fs.readFileSync(
path.join(__dirname, '..', '..', 'payload.config.ts'), 'utf8',
);
expect(config).toMatch(/prodMigrations:\s*migrations/);
});
});
@@ -0,0 +1,45 @@
/**
* Guards the one thing about Payload's mounting that fails silently.
*
* payload.config.ts itself cannot be imported here — Payload is ESM-only and
* next/jest will not transform it — so this asserts the shared constants and
* that the config actually wires them in, by reading its source. The live
* proof that /api still reaches FastAPI is the e2e journeys, which call
* /api/schools against the running app.
*/
import fs from 'fs';
import path from 'path';
import { PAYLOAD_API_ROUTE, PAYLOAD_ADMIN_ROUTE } from '@/lib/payloadRoutes';
const CONFIG = fs.readFileSync(
path.join(__dirname, '..', '..', 'payload.config.ts'),
'utf8',
);
describe('payload mount points', () => {
it('serves the CMS API from /cms-api, never /api', () => {
// /api is the FastAPI proxy's catch-all. Payload's default would be
// swallowed by it and forwarded to the backend, silently.
expect(PAYLOAD_API_ROUTE).toBe('/cms-api');
expect(PAYLOAD_API_ROUTE).not.toBe('/api');
});
it('serves the admin panel from /admin', () => {
expect(PAYLOAD_ADMIN_ROUTE).toBe('/admin');
});
it('wires both constants into the Payload config', () => {
expect(CONFIG).toContain('PAYLOAD_API_ROUTE');
expect(CONFIG).toContain('PAYLOAD_ADMIN_ROUTE');
});
it('never hardcodes a routes block that could drift from the constants', () => {
expect(CONFIG).not.toMatch(/routes:\s*\{[^}]*api:\s*['"]/);
});
it('isolates CMS tables in their own postgres schema', () => {
// Blog content must sit outside `public`, where the app tables, Airflow's
// metadata and scripts/migrate_csv_to_db.py --drop all live.
expect(CONFIG).toMatch(/schemaName:\s*['"]payload['"]/);
});
});
@@ -21,7 +21,7 @@ import {
import { nationalAveragesFixture } from './schoolFixtures'; import { nationalAveragesFixture } from './schoolFixtures';
// The shell calls useComparison(), which throws outside the provider. In the // The shell calls useComparison(), which throws outside the provider. In the
// app this wrapper comes from app/layout.tsx. // app this wrapper comes from app/(frontend)/layout.tsx.
function withProviders(ui: ReactNode) { function withProviders(ui: ReactNode) {
return <ComparisonProvider>{ui}</ComparisonProvider>; return <ComparisonProvider>{ui}</ComparisonProvider>;
} }
@@ -0,0 +1,82 @@
.page {
max-width: 42rem;
margin: 0 auto;
padding: 2.5rem 1.25rem 4rem;
}
.header {
display: flex;
align-items: center;
gap: 1.25rem;
margin-bottom: 2rem;
}
.portrait {
border-radius: 50%;
border: 2px solid var(--border);
object-fit: cover;
flex-shrink: 0;
}
.kicker {
font-family: var(--font-ui);
font-size: 0.75rem;
font-weight: 600;
text-transform: uppercase;
letter-spacing: 0.06em;
color: var(--brand);
margin: 0 0 0.35rem;
}
.heading {
font-family: var(--font-display);
font-size: clamp(1.5rem, 4vw, 2rem);
font-weight: 700;
line-height: 1.2;
color: var(--text-primary);
margin: 0;
}
.subheading {
font-family: var(--font-display);
font-size: 1.15rem;
font-weight: 600;
color: var(--text-primary);
margin: 2.25rem 0 0.75rem;
}
.prose p {
font-family: var(--font-ui);
font-size: 1rem;
line-height: 1.7;
color: var(--text-secondary);
margin: 0 0 1.1rem;
}
/* The opening paragraph carries the page. Larger, and in the primary ink
rather than the secondary, so it reads as a voice rather than as body copy.
Must stay in the descendant form: `.prose p` scores (0,1,1) and would beat a
bare `.lede` at (0,1,0), so simplifying this selector silently reverts the
lede to ordinary body copy. */
.prose .lede {
font-size: 1.125rem;
color: var(--text-primary);
}
.link {
color: var(--brand);
font-weight: 600;
}
.link:hover {
color: var(--brand-strong);
}
@media (max-width: 480px) {
.header {
flex-direction: column;
align-items: flex-start;
gap: 1rem;
}
}
+143
View File
@@ -0,0 +1,143 @@
import type { Metadata } from 'next';
import Image from 'next/image';
import { notFound } from 'next/navigation';
import { absoluteUrl } from '@/lib/site';
import { getFlags } from '@/lib/flags';
import { personJsonLd, organizationJsonLd } from '@/lib/jsonld';
import styles from './About.module.css';
export const metadata: Metadata = {
title: 'About',
description:
'Who builds schoolcompare, why it exists, and where its numbers come from.',
alternates: { canonical: absoluteUrl('/about') },
};
/*
* Gated on about_page. The default 300s read is the right floor here: this
* page declares no revalidate of its own, so nothing is lost by it, and a flip
* lands within five minutes.
*
* notFound(), not a redirect: while the flag is dark this URL does not exist,
* and a 404 is what tells a crawler not to keep it.
*/
export default async function AboutPage() {
const flags = await getFlags();
if (flags.about_page !== true) notFound();
const jsonLd = {
'@context': 'https://schema.org',
'@graph': [personJsonLd(), organizationJsonLd()],
};
return (
<div className={styles.page}>
<script
type="application/ld+json"
dangerouslySetInnerHTML={{ __html: JSON.stringify(jsonLd) }}
/>
<header className={styles.header}>
<Image
src="/brand/tudor.jpg"
alt="Tudor, who builds schoolcompare"
width={96}
height={96}
className={styles.portrait}
priority
/>
<div>
<p className={styles.kicker}>Who&apos;s behind this</p>
<h1 className={styles.heading}>I&apos;m Tudor. I built this site.</h1>
</div>
</header>
<div className={styles.prose}>
<p className={styles.lede}>
I&apos;m a parent in south-west London. When we started looking at
primary schools, I found the information I needed was all published,
and almost impossible to hold in one place.
</p>
<p>
SATs results were in one government table. Ofsted judgements were in a
separate service, in a format that had just changed. Admissions
distances were buried in council PDFs, a different one per borough,
each with its own layout. I ended up building a spreadsheet, and then
I got tired of the spreadsheet.
</p>
<p>
So I built this instead. It pulls the official figures into one place
and puts them side by side, which is what I wanted and could not find.
</p>
<h2 className={styles.subheading}>I&apos;m not an education expert</h2>
<p>
I want to be straightforward about that. I&apos;m not a teacher, a
governor, or an education researcher. I have no qualification that
makes my opinion about a school worth more than yours.
</p>
<p>
What I do have is the problem itself. I&apos;m going through primary
admissions right now, and I work with data for a living. That
combination is enough to take published figures and present them
honestly. It is not enough to tell you which school is right for your
child, and this site never tries to.
</p>
<h2 className={styles.subheading}>Where the numbers come from</h2>
<p>
Everything here is official published data: Key Stage 2 and Key Stage
4 results and school characteristics from the Department for
Education, inspection outcomes from Ofsted, and admissions data from
local authorities. Nothing is estimated, modelled or filled in. Where
a figure is missing, the page says so rather than showing a guess.
</p>
<p>
This is an independent site. It is not affiliated with the Department
for Education or with Ofsted, and nobody pays to appear on it or to
rank higher.
</p>
<h2 className={styles.subheading}>What the data can&apos;t tell you</h2>
<p>
A school is not its results. The figures here describe one year group,
on a handful of days, measured in a way that suits national statistics
rather than your child. A small cohort makes percentages swing wildly.
In a class of thirty, one pupil is worth more than three points.
Results say nothing at all about whether a child will be happy
somewhere.
</p>
<p>
I try to build that honesty into the site rather than just say it
here. Special schools and pupil referral units are never compared
against a mainstream national average, because that comparison is
meaningless and makes good schools look like failing ones. Where a
number is unreliable, the aim is for the page to tell you before you
draw a conclusion from it.
</p>
<h2 className={styles.subheading}>If something&apos;s wrong</h2>
<p>
Tell me and I&apos;ll fix it. If a figure looks wrong, or a page gives
a misleading impression of a school, I genuinely want to know.
It&apos;s the fastest way this gets better.
</p>
<p>
<a href="mailto:contact@schoolcompare.co.uk" className={styles.link}>
contact@schoolcompare.co.uk
</a>
</p>
</div>
</div>
);
}
@@ -26,8 +26,27 @@ function backendBase(): string {
const STRIPPED_RESPONSE_HEADERS = ['content-encoding', 'content-length', 'transfer-encoding', 'connection']; const STRIPPED_RESPONSE_HEADERS = ['content-encoding', 'content-length', 'transfer-encoding', 'connection'];
const METHODS_WITH_BODY = new Set(['POST', 'PUT', 'PATCH', 'DELETE']); const METHODS_WITH_BODY = new Set(['POST', 'PUT', 'PATCH', 'DELETE']);
/*
* API paths this public proxy must not forward.
*
* Matched on the first segment, exactly — a prefix match would take
* /api/flagship down with /api/flags.
*
* `flags` is here because GET /api/flags names every unreleased feature the
* codebase knows about, along with whether it is on. Publishing that defeats
* the point of shipping dark. Next reads it server-side via FASTAPI_URL, on
* the Docker network, which never transits this route.
*
* Anything else internal-only belongs here too.
*/
const INTERNAL_ONLY_SEGMENTS = new Set(['flags']);
async function handler(req: NextRequest, ctx: { params: Promise<{ path: string[] }> }) { async function handler(req: NextRequest, ctx: { params: Promise<{ path: string[] }> }) {
const { path } = await ctx.params; const { path } = await ctx.params;
if (INTERNAL_ONLY_SEGMENTS.has(path[0])) {
return NextResponse.json({ detail: 'Not Found' }, { status: 404 });
}
const target = `${backendBase()}/${path.join('/')}${req.nextUrl.search}`; const target = `${backendBase()}/${path.join('/')}${req.nextUrl.search}`;
const headers = new Headers(req.headers); const headers = new Headers(req.headers);
@@ -0,0 +1,76 @@
.page {
max-width: 42rem;
margin: 0 auto;
padding: 2.5rem 1.25rem 4rem;
}
.header { margin-bottom: 2.5rem; }
.kicker {
font-family: var(--font-ui);
font-size: 0.75rem;
font-weight: 600;
text-transform: uppercase;
letter-spacing: 0.06em;
color: var(--brand);
margin: 0 0 0.35rem;
}
.heading {
font-family: var(--font-display);
font-size: clamp(1.5rem, 4vw, 2rem);
font-weight: 700;
line-height: 1.2;
color: var(--text-primary);
margin: 0 0 0.75rem;
}
.standfirst {
font-family: var(--font-ui);
font-size: 1.05rem;
line-height: 1.65;
color: var(--text-secondary);
margin: 0;
}
.list { list-style: none; padding: 0; margin: 0; }
.item {
padding: 1.5rem 0;
border-top: 1px solid var(--border);
}
.date {
font-family: var(--font-ui);
font-size: 0.8rem;
color: var(--text-muted);
/* Inter's tabular numerals keep a column of dates aligned. */
font-variant-numeric: tabular-nums;
}
.itemTitle {
font-family: var(--font-display);
font-size: 1.25rem;
font-weight: 600;
line-height: 1.3;
margin: 0.35rem 0 0.5rem;
}
.itemLink { color: var(--text-primary); text-decoration: none; }
.itemLink:hover { color: var(--brand); }
.excerpt {
font-family: var(--font-ui);
font-size: 0.95rem;
line-height: 1.65;
color: var(--text-secondary);
margin: 0;
}
.empty {
font-family: var(--font-ui);
color: var(--text-muted);
}
.link { color: var(--brand); font-weight: 600; }
.link:hover { color: var(--brand-strong); }
@@ -0,0 +1,87 @@
.page {
max-width: 42rem;
margin: 0 auto;
padding: 2.5rem 1.25rem 4rem;
}
.crumb {
font-family: var(--font-ui);
font-size: 0.85rem;
margin-bottom: 1.25rem;
}
.heading {
font-family: var(--font-display);
font-size: clamp(1.6rem, 5vw, 2.25rem);
font-weight: 700;
line-height: 1.2;
color: var(--text-primary);
margin: 0 0 0.75rem;
}
.byline {
font-family: var(--font-ui);
font-size: 0.9rem;
color: var(--text-muted);
margin: 0 0 2rem;
}
.hero {
width: 100%;
height: auto;
border-radius: 10px;
border: 1px solid var(--border);
margin-bottom: 2rem;
}
/* Rich-text output: the editor emits plain elements, so these are styled by
descendant selector rather than by class. */
.prose p {
font-family: var(--font-ui);
font-size: 1rem;
line-height: 1.7;
color: var(--text-secondary);
margin: 0 0 1.1rem;
}
.prose h2 {
font-family: var(--font-display);
font-size: 1.25rem;
font-weight: 600;
color: var(--text-primary);
margin: 2.25rem 0 0.75rem;
}
.prose h3 {
font-family: var(--font-display);
font-size: 1.05rem;
font-weight: 600;
color: var(--text-primary);
margin: 1.75rem 0 0.6rem;
}
.prose ul,
.prose ol {
font-family: var(--font-ui);
font-size: 1rem;
line-height: 1.7;
color: var(--text-secondary);
padding-left: 1.35rem;
margin: 0 0 1.1rem;
}
.prose li { margin-bottom: 0.4rem; }
.prose a { color: var(--brand); font-weight: 500; }
.prose a:hover { color: var(--brand-strong); }
.prose blockquote {
border-left: 3px solid var(--border-strong);
padding-left: 1rem;
margin: 1.5rem 0;
color: var(--text-muted);
font-style: italic;
}
.link { color: var(--brand); font-weight: 600; }
.link:hover { color: var(--brand-strong); }
@@ -0,0 +1,194 @@
import { cache } from 'react';
import type { Metadata } from 'next';
import Link from 'next/link';
import { notFound } from 'next/navigation';
import { RichText } from '@payloadcms/richtext-lexical/react';
import type { JSXConvertersFunction } from '@payloadcms/richtext-lexical/react';
import { getCachedPayload } from '@/lib/payload';
import type { Post, Media } from '@/payload-types';
import { absoluteUrl } from '@/lib/site';
import { getFlags } from '@/lib/flags';
import {
blogPostingJsonLd,
breadcrumbJsonLd,
personJsonLd,
organizationJsonLd,
} from '@/lib/jsonld';
import { CalloutBlock } from '@/components/blog/CalloutBlock';
import styles from './Post.module.css';
/*
* ISR. Unlike the index, this route has a dynamic param and no
* generateStaticParams, so there is nothing for the build to prerender: each
* post is generated on first request and cached until the collection's
* afterChange hook revalidates it. That hook is what makes an edit to an
* already-published post appear immediately.
*/
export const revalidate = 3600;
/**
* heroImage is `number | Media | null`: an id when the query is shallow, the
* populated document at depth 1. Both pages query at depth 1, but narrowing
* rather than asserting keeps it correct if that ever changes.
*/
function heroOf(post: Post): Media | null {
return typeof post.heroImage === 'object' && post.heroImage !== null
? post.heroImage
: null;
}
/**
* Spreads the default converters and adds the one custom block.
*
* Without the spread, every default node type — paragraphs, headings, links —
* loses its renderer and the post body comes out empty.
*/
const calloutConverters: JSXConvertersFunction = ({ defaultConverters }) => ({
...defaultConverters,
blocks: {
// Annotated because the generic block converter cannot infer a custom
// block's field shape; String() guards the values regardless.
callout: ({ node }: { node: { fields: Record<string, unknown> } }) => (
<CalloutBlock
tone={String(node.fields.tone ?? 'caveat')}
body={String(node.fields.body ?? '')}
/>
),
},
});
/**
* Wrapped in React's cache() because Next calls generateMetadata and the page
* component separately for the same request — without it, every post view runs
* this query against Postgres twice. cache() dedupes within a single request
* only, so it never serves one visitor's request from another's.
*/
const findPost = cache(async (slug: string) => {
const payload = await getCachedPayload();
const { docs } = await payload.find({
collection: 'posts',
where: { slug: { equals: slug }, _status: { equals: 'published' } },
limit: 1,
depth: 1,
});
return docs[0] ?? null;
});
function summarise(post: Post) {
return {
title: post.title,
slug: post.slug,
excerpt: post.excerpt,
publishedAt: post.publishedAt,
};
}
export async function generateMetadata(
{ params }: { params: Promise<{ slug: string }> },
): Promise<Metadata> {
const { slug } = await params;
const post = await findPost(slug);
if (!post) return { title: 'Not found' };
const hero = heroOf(post);
return {
title: post.title,
description: post.excerpt,
alternates: { canonical: absoluteUrl(`/blog/${post.slug}`) },
openGraph: {
type: 'article',
title: post.title,
description: post.excerpt,
url: absoluteUrl(`/blog/${post.slug}`),
publishedTime: post.publishedAt,
// A post with a hero image shares that; one without falls through to the
// generated share card at app/opengraph-image.tsx.
...(hero?.url ? { images: [{ url: hero.url }] } : {}),
},
};
}
export default async function PostPage(
{ params }: { params: Promise<{ slug: string }> },
) {
const { slug } = await params;
/*
* Flags read at this route's own declared floor, so gating costs it nothing.
* Checked before the post is fetched: a dark blog should not query Payload.
*/
const flags = await getFlags(3600);
if (flags.blog !== true) notFound();
const namedAuthor = flags.about_page === true;
const post = await findPost(slug);
if (!post) notFound();
const summary = summarise(post);
const hero = heroOf(post);
const jsonLd = {
'@context': 'https://schema.org',
/*
* The Person entity is anchored at /about#tudor, so it is declared only
* when that page exists. Claiming an author whose URL 404s is a worse
* signal than attributing the post to the publisher.
*/
'@graph': [
blogPostingJsonLd(summary, { namedAuthor }),
breadcrumbJsonLd(summary),
...(namedAuthor ? [personJsonLd()] : []),
organizationJsonLd(),
],
};
return (
<article className={styles.page}>
<script
type="application/ld+json"
dangerouslySetInnerHTML={{ __html: JSON.stringify(jsonLd) }}
/>
<nav className={styles.crumb}>
<Link href="/blog" className={styles.link}>Blog</Link>
</nav>
<h1 className={styles.heading}>{summary.title}</h1>
<p className={styles.byline}>
{/* Unlinked while about_page is dark; the flags are independent. */}
By {namedAuthor
? <Link href="/about" className={styles.link}>Tudor</Link>
: 'Tudor'}
{' · '}
<time dateTime={summary.publishedAt}>
{new Date(summary.publishedAt).toLocaleDateString('en-GB', {
day: 'numeric',
month: 'long',
year: 'numeric',
})}
</time>
</p>
{/*
A plain <img>, not next/image: Payload already generated the sized
derivatives on upload (Media's imageSizes), so routing it through the
optimizer would resize an image that is already the right size.
*/}
{hero?.url && (
<img
className={styles.hero}
src={hero.url}
alt={hero.alt ?? ''}
width={hero.width ?? undefined}
height={hero.height ?? undefined}
/>
)}
<div className={styles.prose}>
<RichText data={post.content} converters={calloutConverters} />
</div>
</article>
);
}
+85
View File
@@ -0,0 +1,85 @@
import type { Metadata } from 'next';
import Link from 'next/link';
import { notFound } from 'next/navigation';
import { getCachedPayload } from '@/lib/payload';
import { absoluteUrl } from '@/lib/site';
import { getFlags } from '@/lib/flags';
import styles from './Blog.module.css';
/*
* Dynamic, not ISR.
*
* This route has no dynamic params, so Next prerenders it at build time — and
* CI builds the image with no database reachable, which fails the build. It is
* a single indexed query against Postgres on the same Docker network, so
* rendering per request is cheap, and it means a newly published post appears
* here immediately rather than waiting on a revalidation.
*/
export const dynamic = 'force-dynamic';
export const metadata: Metadata = {
title: 'Blog',
description:
'Notes on what school performance data shows, and what it does not.',
alternates: { canonical: absoluteUrl('/blog') },
};
function formatDate(value: string) {
return new Date(value).toLocaleDateString('en-GB', {
day: 'numeric',
month: 'long',
year: 'numeric',
});
}
export default async function BlogIndexPage() {
const flags = await getFlags();
if (flags.blog !== true) notFound();
const payload = await getCachedPayload();
const { docs } = await payload.find({
collection: 'posts',
where: { _status: { equals: 'published' } },
sort: '-publishedAt',
limit: 50,
depth: 0,
});
return (
<div className={styles.page}>
<header className={styles.header}>
<p className={styles.kicker}>Blog</p>
<h1 className={styles.heading}>Notes on the numbers</h1>
<p className={styles.standfirst}>
What school performance data shows, what it doesn&apos;t, and how to
read it without being misled. Written by{' '}
{/* Plain text when about_page is dark: the two flags are
independent, so this link would otherwise point at a 404. */}
{flags.about_page === true
? <Link href="/about" className={styles.link}>Tudor</Link>
: 'Tudor'}.
</p>
</header>
{docs.length === 0 ? (
<p className={styles.empty}>No posts yet.</p>
) : (
<ul className={styles.list}>
{docs.map((post) => (
<li key={post.id} className={styles.item}>
<time className={styles.date} dateTime={String(post.publishedAt)}>
{formatDate(String(post.publishedAt))}
</time>
<h2 className={styles.itemTitle}>
<Link href={`/blog/${post.slug}`} className={styles.itemLink}>
{post.title}
</Link>
</h2>
<p className={styles.excerpt}>{post.excerpt}</p>
</li>
))}
</ul>
)}
</div>
);
}
@@ -0,0 +1,58 @@
import { getCachedPayload } from '@/lib/payload';
import { absoluteUrl } from '@/lib/site';
import { getFlags } from '@/lib/flags';
/*
* Dynamic, not ISR.
*
* This route has no dynamic params, so Next prerenders it at build time — and
* CI builds the image with no database reachable, which fails the build. It is
* a single indexed query against Postgres on the same Docker network, so
* rendering per request is cheap, and it means a newly published post appears
* here immediately rather than waiting on a revalidation.
*/
export const dynamic = 'force-dynamic';
function escapeXml(value: string): string {
return value.replace(/[<>&'"]/g, (char) =>
({ '<': '&lt;', '>': '&gt;', '&': '&amp;', "'": '&apos;', '"': '&quot;' }[char]!));
}
export async function GET() {
// A dark blog has no feed. 404 rather than an empty channel: an empty feed
// is a live feed with nothing in it, which a reader would keep polling.
const flags = await getFlags();
if (flags.blog !== true) return new Response('Not found', { status: 404 });
const payload = await getCachedPayload();
const { docs } = await payload.find({
collection: 'posts',
where: { _status: { equals: 'published' } },
sort: '-publishedAt',
limit: 50,
depth: 0,
});
const items = docs.map((post) => `
<item>
<title>${escapeXml(String(post.title))}</title>
<link>${absoluteUrl(`/blog/${post.slug}`)}</link>
<guid isPermaLink="true">${absoluteUrl(`/blog/${post.slug}`)}</guid>
<description>${escapeXml(String(post.excerpt))}</description>
<pubDate>${new Date(String(post.publishedAt)).toUTCString()}</pubDate>
</item>`).join('');
const xml = `<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0">
<channel>
<title>schoolcompare blog</title>
<link>${absoluteUrl('/blog')}</link>
<description>What school performance data shows, and what it does not.</description>
<language>en-GB</language>${items}
</channel>
</rss>`;
return new Response(xml, {
headers: { 'Content-Type': 'application/rss+xml; charset=utf-8' },
});
}
File renamed without changes.
@@ -0,0 +1,68 @@
/*
* A second sitemap for the URLs Next owns.
*
* /sitemap.xml is proxied from FastAPI (app/(frontend)/sitemap.xml), which
* knows nothing about Payload — the backend and frontend ship as separate
* images. Rather than teach it, the Next-owned URLs get their own sitemap and
* robots.txt lists both.
*/
import { getCachedPayload } from '@/lib/payload';
import { absoluteUrl } from '@/lib/site';
import { getFlags } from '@/lib/flags';
/*
* Dynamic, not ISR.
*
* This route has no dynamic params, so Next prerenders it at build time — and
* CI builds the image with no database reachable, which fails the build. It is
* a single indexed query against Postgres on the same Docker network, so
* rendering per request is cheap, and it means a newly published post appears
* here immediately rather than waiting on a revalidation.
*/
export const dynamic = 'force-dynamic';
export async function GET() {
/*
* A dark page must not be advertised. Submitting a URL that 404s is the one
* thing a sitemap is not allowed to do, so each entry is gated on the same
* flag that gates the page itself.
*
* With both flags dark this emits a valid, empty <urlset> rather than a 404:
* robots.txt names this sitemap unconditionally, and an empty sitemap is a
* well-formed statement that there is nothing here yet.
*/
const flags = await getFlags();
const aboutEnabled = flags.about_page === true;
const blogEnabled = flags.blog === true;
// Only query Payload when the blog is actually being advertised.
const docs = blogEnabled
? (await (await getCachedPayload()).find({
collection: 'posts',
where: { _status: { equals: 'published' } },
sort: '-publishedAt',
limit: 500,
depth: 0,
})).docs
: [];
const urls: Array<{ loc: string; lastmod: string | null }> = [
...(aboutEnabled ? [{ loc: absoluteUrl('/about'), lastmod: null }] : []),
...(blogEnabled ? [{ loc: absoluteUrl('/blog'), lastmod: null }] : []),
...docs.map((post) => ({
loc: absoluteUrl(`/blog/${post.slug}`),
lastmod: new Date(String(post.updatedAt ?? post.publishedAt)).toISOString(),
})),
];
const xml = `<?xml version="1.0" encoding="UTF-8"?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
${urls.map(({ loc, lastmod }) =>
` <url><loc>${loc}</loc>${lastmod ? `<lastmod>${lastmod}</lastmod>` : ''}</url>`,
).join('\n')}
</urlset>`;
return new Response(xml, {
headers: { 'Content-Type': 'application/xml; charset=utf-8' },
});
}
@@ -28,6 +28,9 @@
--bg-primary: #FAFAF8; /* Warm White */ --bg-primary: #FAFAF8; /* Warm White */
--bg-secondary: #F5EFE6; /* Sand — hero panels, sunken rows */ --bg-secondary: #F5EFE6; /* Sand — hero panels, sunken rows */
--bg-card: #FFFFFF; --bg-card: #FFFFFF;
/* For gradients that have to fade to the card colour. A hardcoded white
ramp reads as a bright band against a dark card. */
--bg-card-rgb: 255, 255, 255;
--surface-inverse: #0F766E; --surface-inverse: #0F766E;
/* ── Ink ────────────────────────────────────────────────────────── */ /* ── Ink ────────────────────────────────────────────────────────── */
@@ -102,6 +105,23 @@
--series-7: #0E7A86; --series-7: #0E7A86;
--series-8: #8A4A6B; --series-8: #8A4A6B;
/* ── Destination measures ───────────────────────────────────────────
Education is one hue in three steps (school-like -> college-like) so the
education destinations read as one family; apprenticeship and employment
are separate hues. The absence is neutral and HATCHED, never a colour:
"activity not captured" covers independent schools, moving abroad and
training DfE holds no data on, so rendering it as a bad outcome would be
a factual error. The hatch is also the secondary encoding that rescues
the neutral/blue pair, which separates at only dE 7.6 as flat fills.
Every other adjacent pair clears dE 10.9 under protanopia. */
--dest-sixthform: #0F766E;
--dest-sfcollege: #4A9E96;
--dest-fecollege: #7CBFB8;
--dest-apprentice: #806200;
--dest-employment: #2F6F8F;
--dest-none: #6B7580;
--dest-none-hatch: rgba(107, 117, 128, 0.34);
/* ── Phase: category, desaturated so it stays under the status hues ── */ /* ── Phase: category, desaturated so it stays under the status hues ── */
--phase-primary: #0F766E; --phase-primary: #0F766E;
--phase-primary-bg: rgba(167, 215, 197, 0.40); --phase-primary-bg: rgba(167, 215, 197, 0.40);
@@ -234,6 +254,7 @@
--bg-primary: #111A20; --bg-primary: #111A20;
--bg-secondary: #16222A; --bg-secondary: #16222A;
--bg-card: #18242C; --bg-card: #18242C;
--bg-card-rgb: 24, 36, 44;
--surface-inverse: #E9EEF0; --surface-inverse: #E9EEF0;
--text-primary: #E9EEF0; --text-primary: #E9EEF0;
@@ -289,6 +310,17 @@
--series-7: #6FD0DC; --series-7: #6FD0DC;
--series-8: #D99BB8; --series-8: #D99BB8;
/* Destinations. Not a naive inversion: the education ramp reverses
direction so its darkest step stays the one furthest from the
school, and each step is re-checked against the dark card. */
--dest-sixthform: #5FC7BB;
--dest-sfcollege: #3E9B92;
--dest-fecollege: #2A716B;
--dest-apprentice: #EFC658;
--dest-employment: #8FB4D9;
--dest-none: #8B9AA1;
--dest-none-hatch: rgba(139, 154, 161, 0.34);
--phase-primary: #5FC7BB; --phase-primary: #5FC7BB;
--phase-primary-bg: rgba(95, 199, 187, 0.16); --phase-primary-bg: rgba(95, 199, 187, 0.16);
--phase-primary-text: #8ADACF; --phase-primary-text: #8ADACF;
@@ -584,6 +616,35 @@ html .leaflet-bar a:hover {
color: var(--text-primary); color: var(--text-primary);
} }
/*
* The popup, which leaflet.css paints `background: white; color: #333` on both
* the card and its tip. The content LeafletMapInner binds into it is themed —
* the school name and the headline figure are `var(--text-primary)` — so in
* dark mode that was #E9EEF0 on #FFFFFF, a contrast ratio of 1.17:1. The name
* and the number were the two least readable things on the page.
*
* Moving the surface onto --bg-card fixes every foreground at once rather than
* one at a time: the muted phase line goes 2.90:1 -> 5.45:1, the vs-national
* delta 1.94:1 -> 8.14:1, the Ofsted badge 1.74:1 -> 9.11:1. In light mode
* --bg-card is #FFFFFF, so the popup looks as it always did.
*/
html .leaflet-popup-content-wrapper,
html .leaflet-popup-tip {
background: var(--bg-card);
color: var(--text-primary);
}
/* Leaflet's own selector is `.leaflet-container a.leaflet-popup-close-button`
at 0,2,1 — an `html` prefix alone would lose to it. */
html .leaflet-container a.leaflet-popup-close-button {
color: var(--text-muted);
}
html .leaflet-container a.leaflet-popup-close-button:hover,
html .leaflet-container a.leaflet-popup-close-button:focus {
color: var(--text-primary);
}
/* Main content column */ /* Main content column */
.main { .main {
max-width: 1400px; max-width: 1400px;
@@ -4,8 +4,10 @@ import Script from 'next/script';
import { Navigation } from '@/components/Navigation'; import { Navigation } from '@/components/Navigation';
import { Footer } from '@/components/Footer'; import { Footer } from '@/components/Footer';
import { ComparisonToast } from '@/components/ComparisonToast'; import { ComparisonToast } from '@/components/ComparisonToast';
import { RouteTrail } from '@/components/RouteTrail';
import { ComparisonProvider } from '@/context/ComparisonProvider'; import { ComparisonProvider } from '@/context/ComparisonProvider';
import { SITE_URL } from '@/lib/site'; import { SITE_URL } from '@/lib/site';
import { getFlags } from '@/lib/flags';
import './globals.css'; import './globals.css';
// Manrope carries headings and key messaging — the guideline's "friendly, // Manrope carries headings and key messaging — the guideline's "friendly,
@@ -77,11 +79,28 @@ export const metadata: Metadata = {
}, },
}; };
export default function RootLayout({ /*
* The footer's About and Blog links are flagged, which makes this the one
* place on the site that reads a flag on every route.
*
* 604800 is deliberate and load-bearing: it is the revalidate every SEO route
* here already declares. Next pins a route to the LOWEST revalidate among its
* fetches, so reading flags at the 300s default would drop the whole school
* and place corpus from a weekly cache to a 5-minute one — a large origin-load
* regression to hide two footer links.
*
* The cost is latency in one direction only. The pages themselves read the
* same flags at their own floors and flip within minutes; the footer links
* follow within a week. Turning a feature on early therefore shows the page
* before its footer link, which is harmless. Turning one off leaves a link to
* a 404 until the cache turns over, so a rollback that matters wants a purge.
*/
export default async function RootLayout({
children, children,
}: Readonly<{ }: Readonly<{
children: React.ReactNode; children: React.ReactNode;
}>) { }>) {
const flags = await getFlags(604800);
return ( return (
// The font variable classes must sit on <html>, not <body>. globals.css // The font variable classes must sit on <html>, not <body>. globals.css
// declares --font-display on :root as var(--font-manrope) and --font-ui as // declares --font-display on :root as var(--font-manrope) and --font-ui as
@@ -114,6 +133,10 @@ export default function RootLayout({
/> />
</head> </head>
<body> <body>
{/* Records every route so funnel attribution has a previous page to
name. document.referrer cannot: a soft navigation creates no
document, so the browser never updates it. */}
<RouteTrail />
<ComparisonProvider> <ComparisonProvider>
<a href="#main-content" className="skip-link">Skip to main content</a> <a href="#main-content" className="skip-link">Skip to main content</a>
<Navigation /> <Navigation />
@@ -121,7 +144,10 @@ export default function RootLayout({
{children} {children}
</main> </main>
<ComparisonToast /> <ComparisonToast />
<Footer /> <Footer
aboutEnabled={flags.about_page === true}
blogEnabled={flags.blog === true}
/>
</ComparisonProvider> </ComparisonProvider>
</body> </body>
</html> </html>
@@ -8,6 +8,7 @@ import type { Metadata } from 'next';
import { fetchSchools, fetchFilters, fetchDataInfo } from '@/lib/api'; import { fetchSchools, fetchFilters, fetchDataInfo } from '@/lib/api';
import { formatAcademicYear } from '@/lib/utils'; import { formatAcademicYear } from '@/lib/utils';
import { HomeView } from '@/components/HomeView'; import { HomeView } from '@/components/HomeView';
import { getFlags } from '@/lib/flags';
import { HowItWorksSection } from '@/components/HowItWorksSection'; import { HowItWorksSection } from '@/components/HowItWorksSection';
import { EditorialSection } from '@/components/EditorialSection'; import { EditorialSection } from '@/components/EditorialSection';
@@ -63,6 +64,11 @@ export default async function HomePage({ searchParams }: HomePageProps) {
// Await search params (Next.js 15 requirement) // Await search params (Next.js 15 requirement)
const params = await searchParams; const params = await searchParams;
// Server-read: no flag value reaches the browser bundle. Threaded down to
// both FilterBar instances via HomeView.
const flags = await getFlags();
const autosuggest = flags.school_autosuggest === true;
// Parse search params // Parse search params
const page = parseInt(params.page || '1'); const page = parseInt(params.page || '1');
const radius = params.radius ? parseFloat(params.radius) : undefined; const radius = params.radius ? parseFloat(params.radius) : undefined;
@@ -111,6 +117,7 @@ export default async function HomePage({ searchParams }: HomePageProps) {
const years = dataInfo?.years_available ?? []; const years = dataInfo?.years_available ?? [];
return ( return (
<HomeView <HomeView
autosuggest={autosuggest}
initialSchools={schoolsData} initialSchools={schoolsData}
filters={resolvedFilters} filters={resolvedFilters}
totalSchools={total} totalSchools={total}
@@ -131,6 +138,7 @@ export default async function HomePage({ searchParams }: HomePageProps) {
const emptyFilters = { local_authorities: [], school_types: [], years: [], phases: [], genders: [], admissions_policies: [] }; const emptyFilters = { local_authorities: [], school_types: [], years: [], phases: [], genders: [], admissions_policies: [] };
return ( return (
<HomeView <HomeView
autosuggest={autosuggest}
initialSchools={{ schools: [], page: 1, page_size: 50, total: 0, total_pages: 0 }} initialSchools={{ schools: [], page: 1, page_size: 50, total: 0, total_pages: 0 }}
filters={emptyFilters} filters={emptyFilters}
totalSchools={null} totalSchools={null}
File renamed without changes.
@@ -148,7 +148,7 @@ export default async function SchoolPage({ params }: SchoolPageProps) {
notFound(); notFound();
} }
const { school_info, yearly_data, absence_data, ofsted, census, admissions, admissions_history, admission_distance, deprivation, finance } = data; const { school_info, yearly_data, absence_data, ofsted, census, admissions, admissions_history, admission_distance, deprivation, finance, destinations } = data;
// Redirect bare URN to canonical slug URL // Redirect bare URN to canonical slug URL
const canonicalSlug = schoolUrl(urn, school_info.school_name).replace('/school/', ''); const canonicalSlug = schoolUrl(urn, school_info.school_name).replace('/school/', '');
@@ -171,6 +171,7 @@ export default async function SchoolPage({ params }: SchoolPageProps) {
schoolInfo: school_info, yearlyData: yearly_data, schoolInfo: school_info, yearlyData: yearly_data,
absenceData: absence_data, census: census ?? null, absenceData: absence_data, census: census ?? null,
deprivation: deprivation ?? null, finance: finance ?? null, deprivation: deprivation ?? null, finance: finance ?? null,
destinations: destinations ?? null,
}; };
const primaryFlags = computeSchoolFlags(sectionInput); const primaryFlags = computeSchoolFlags(sectionInput);
const secondaryFlags = computeSecondaryFlags(sectionInput); const secondaryFlags = computeSecondaryFlags(sectionInput);
@@ -232,10 +233,11 @@ export default async function SchoolPage({ params }: SchoolPageProps) {
census={census ?? null} census={census ?? null}
admissions={admissions ?? null} admissions={admissions ?? null}
admissionsHistory={admissions_history ?? []} admissionsHistory={admissions_history ?? []}
admissionDistance={admission_distance ?? null} admissionDistance={admission_distance}
deprivation={deprivation ?? null} deprivation={deprivation ?? null}
finance={finance ?? null} finance={finance ?? null}
nationalAvg={nationalAvg} nationalAvg={nationalAvg}
destinations={destinations ?? null}
flags={secondaryFlags} flags={secondaryFlags}
/> />
</SchoolDetailShell> </SchoolDetailShell>
@@ -37,10 +37,12 @@ export async function generateMetadata({ params }: Props): Promise<Metadata> {
const word = phase === 'secondary' ? 'Secondary' : 'Primary'; const word = phase === 'secondary' ? 'Secondary' : 'Primary';
const { name } = detail.place; const { name } = detail.place;
return { return {
title: { absolute: `${word} Schools in ${name} — Ranked | schoolcompare` }, // Not "Ranked": the table is alphabetical, so the word would be a claim
// the page does not keep.
title: { absolute: `${word} Schools in ${name} | schoolcompare` },
description: description:
`Every ${phase} school in ${name} ranked by results, with Ofsted grades and ` `Every ${phase} school in ${name}, with results, Ofsted grades and the local `
+ `the local average against England.`, + `average against England.`,
alternates: { canonical: absoluteUrl(`/schools/${slug}/${phase}`) }, alternates: { canonical: absoluteUrl(`/schools/${slug}/${phase}`) },
}; };
} }
@@ -58,8 +58,8 @@ export async function generateMetadata({ params }: Props): Promise<Metadata> {
// place title read '... | schoolcompare | schoolcompare'. // place title read '... | schoolcompare | schoolcompare'.
title: { absolute: `Schools in ${name} — Compare ${count} Schools | schoolcompare` }, title: { absolute: `Schools in ${name} — Compare ${count} Schools | schoolcompare` },
description: description:
`Every school in ${name} ranked by SATs and GCSE results, with Ofsted grades, ` `Every school in ${name}, with SATs and GCSE results, Ofsted grades, the local `
+ `the local average against England, and how close you had to live to get a place.`, + `average against England, and how close you had to live to get a place.`,
alternates: { canonical: absoluteUrl(`/schools/${slug}`) }, alternates: { canonical: absoluteUrl(`/schools/${slug}`) },
}; };
} }
@@ -74,9 +74,14 @@ export default async function PlacePage({ params }: Props) {
// defers to its authority rather than publishing a thin page. // defers to its authority rather than publishing a thin page.
if (detail.averages.rwm_expected_pct == null if (detail.averages.rwm_expected_pct == null
&& detail.averages.attainment_8_score == null) { && detail.averages.attainment_8_score == null) {
if (detail.place.parent_authority) { // The API's own slug, which is null when that authority is itself under
redirect(`/schools/authority/${authoritySlug(detail.place.parent_authority)}`); // the threshold and has no page. Re-slugifying the name here would send
} // the reader to a 404 instead of telling them the place has no page.
const target = detail.place.authorities?.[0]?.slug
?? (detail.place.parent_authority
? authoritySlug(detail.place.parent_authority)
: null);
if (target) redirect(`/schools/authority/${target}`);
notFound(); notFound();
} }
@@ -0,0 +1,65 @@
/**
* Phase variants of an authority page.
*
* The spec called for these; the plan built the bare authority route and
* dropped them. Nothing caught it, because the sitemap is written from the
* place registry — which was right about them all along — while the routes
* were written by hand. 302 authority phase URLs were submitted to Google and
* every one 404'd, and every authority page linked to a phase page in the
* *town* namespace, which is a different set of schools entirely.
*
* "Primary schools in Kent" is the query these serve, and it is a real one:
* admissions are authority-run, so the authority is the unit a parent thinks
* in when they have not settled on a town.
*/
import { notFound } from 'next/navigation';
import type { Metadata } from 'next';
import { fetchPlace } from '@/lib/places';
import { fetchNationalAverages } from '@/lib/api';
import { PlaceView } from '@/components/places/PlaceView';
import { absoluteUrl } from '@/lib/site';
interface Props { params: Promise<{ la: string; phase: string }> }
export const revalidate = 604800;
export const dynamicParams = true;
const PHASES = ['primary', 'secondary'] as const;
type Phase = (typeof PHASES)[number];
const isPhase = (v: string): v is Phase => (PHASES as readonly string[]).includes(v);
export async function generateMetadata({ params }: Props): Promise<Metadata> {
const { la, phase } = await params;
if (!isPhase(phase)) return { title: 'Place Not Found' };
const detail = await fetchPlace('authority', la, phase);
if (!detail || detail.schools.length === 0) return { title: 'Place Not Found' };
const word = phase === 'secondary' ? 'Secondary' : 'Primary';
const { name } = detail.place;
return {
// "Local Authority" stays in the title for the same reason it is on the
// bare authority page: 67 town names collide with an authority name, and
// a reader landing on both needs to know which set each covers.
title: { absolute: `${word} Schools in ${name} — Local Authority | schoolcompare` },
description:
`Every ${phase} school in the ${name} local authority, with results, Ofsted `
+ `grades and the authority average against England.`,
alternates: { canonical: absoluteUrl(`/schools/authority/${la}/${phase}`) },
};
}
export default async function AuthorityPhasePage({ params }: Props) {
const { la, phase } = await params;
if (!isPhase(phase)) notFound();
const detail = await fetchPlace('authority', la, phase);
if (!detail || detail.schools.length === 0) notFound();
const national = await fetchNationalAverages().catch(() => null);
const englandAverage = phase === 'secondary'
? national?.secondary?.attainment_8_score ?? null
: national?.primary?.rwm_expected_pct ?? null;
return <PlaceView detail={detail} phase={phase}
englandAverage={englandAverage} neighbours={[]} />;
}
@@ -45,8 +45,8 @@ export async function generateMetadata({ params }: Props): Promise<Metadata> {
return { return {
title: { absolute: `Schools in ${name} — Local Authority | schoolcompare` }, title: { absolute: `Schools in ${name} — Local Authority | schoolcompare` },
description: description:
`All ${count} schools in the ${name} local authority, ranked by SATs and GCSE ` `All ${count} schools in the ${name} local authority, with SATs and GCSE results, `
+ `results, with Ofsted grades and the authority average against England.`, + `Ofsted grades and the authority average against England.`,
alternates: { canonical: absoluteUrl(`/schools/authority/${la}`) }, alternates: { canonical: absoluteUrl(`/schools/authority/${la}`) },
}; };
} }
@@ -40,8 +40,8 @@ export async function generateMetadata({ params }: Props): Promise<Metadata> {
return { return {
title: { absolute: `Schools near ${name} | schoolcompare` }, title: { absolute: `Schools near ${name} | schoolcompare` },
description: description:
`${count} schools in the ${name} postcode district, ranked by results, with ` `${count} schools in the ${name} postcode district, with results, Ofsted grades `
+ `Ofsted grades and how close you had to live to get a place.`, + `and how close you had to live to get a place.`,
alternates: { canonical: absoluteUrl(`/schools/near/${outcode}`) }, alternates: { canonical: absoluteUrl(`/schools/near/${outcode}`) },
}; };
} }
@@ -0,0 +1,16 @@
import type { Metadata } from 'next';
import config from '@payload-config';
import { NotFoundPage, generatePageMetadata } from '@payloadcms/next/views';
import { importMap } from '../importMap.js';
type Args = {
params: Promise<{ segments: string[] }>;
searchParams: Promise<{ [key: string]: string | string[] }>;
};
export const generateMetadata = ({ params, searchParams }: Args): Promise<Metadata> =>
generatePageMetadata({ config, params, searchParams });
export default function NotFound({ params, searchParams }: Args) {
return NotFoundPage({ config, importMap, params, searchParams });
}
@@ -0,0 +1,16 @@
import type { Metadata } from 'next';
import config from '@payload-config';
import { RootPage, generatePageMetadata } from '@payloadcms/next/views';
import { importMap } from '../importMap.js';
type Args = {
params: Promise<{ segments: string[] }>;
searchParams: Promise<{ [key: string]: string | string[] }>;
};
export const generateMetadata = ({ params, searchParams }: Args): Promise<Metadata> =>
generatePageMetadata({ config, params, searchParams });
export default function Page({ params, searchParams }: Args) {
return RootPage({ config, importMap, params, searchParams });
}
@@ -0,0 +1,54 @@
import { RscEntryLexicalCell as RscEntryLexicalCell_44fe37237e0ebf4470c9990d8cb7b07e } from '@payloadcms/richtext-lexical/rsc'
import { RscEntryLexicalField as RscEntryLexicalField_44fe37237e0ebf4470c9990d8cb7b07e } from '@payloadcms/richtext-lexical/rsc'
import { LexicalDiffComponent as LexicalDiffComponent_44fe37237e0ebf4470c9990d8cb7b07e } from '@payloadcms/richtext-lexical/rsc'
import { BlocksFeatureClient as BlocksFeatureClient_e70f5e05f09f93e00b997edb1ef0c864 } from '@payloadcms/richtext-lexical/client'
import { BoldFeatureClient as BoldFeatureClient_e70f5e05f09f93e00b997edb1ef0c864 } from '@payloadcms/richtext-lexical/client'
import { ItalicFeatureClient as ItalicFeatureClient_e70f5e05f09f93e00b997edb1ef0c864 } from '@payloadcms/richtext-lexical/client'
import { UnderlineFeatureClient as UnderlineFeatureClient_e70f5e05f09f93e00b997edb1ef0c864 } from '@payloadcms/richtext-lexical/client'
import { StrikethroughFeatureClient as StrikethroughFeatureClient_e70f5e05f09f93e00b997edb1ef0c864 } from '@payloadcms/richtext-lexical/client'
import { SubscriptFeatureClient as SubscriptFeatureClient_e70f5e05f09f93e00b997edb1ef0c864 } from '@payloadcms/richtext-lexical/client'
import { SuperscriptFeatureClient as SuperscriptFeatureClient_e70f5e05f09f93e00b997edb1ef0c864 } from '@payloadcms/richtext-lexical/client'
import { InlineCodeFeatureClient as InlineCodeFeatureClient_e70f5e05f09f93e00b997edb1ef0c864 } from '@payloadcms/richtext-lexical/client'
import { ParagraphFeatureClient as ParagraphFeatureClient_e70f5e05f09f93e00b997edb1ef0c864 } from '@payloadcms/richtext-lexical/client'
import { HeadingFeatureClient as HeadingFeatureClient_e70f5e05f09f93e00b997edb1ef0c864 } from '@payloadcms/richtext-lexical/client'
import { AlignFeatureClient as AlignFeatureClient_e70f5e05f09f93e00b997edb1ef0c864 } from '@payloadcms/richtext-lexical/client'
import { IndentFeatureClient as IndentFeatureClient_e70f5e05f09f93e00b997edb1ef0c864 } from '@payloadcms/richtext-lexical/client'
import { UnorderedListFeatureClient as UnorderedListFeatureClient_e70f5e05f09f93e00b997edb1ef0c864 } from '@payloadcms/richtext-lexical/client'
import { OrderedListFeatureClient as OrderedListFeatureClient_e70f5e05f09f93e00b997edb1ef0c864 } from '@payloadcms/richtext-lexical/client'
import { ChecklistFeatureClient as ChecklistFeatureClient_e70f5e05f09f93e00b997edb1ef0c864 } from '@payloadcms/richtext-lexical/client'
import { LinkFeatureClient as LinkFeatureClient_e70f5e05f09f93e00b997edb1ef0c864 } from '@payloadcms/richtext-lexical/client'
import { RelationshipFeatureClient as RelationshipFeatureClient_e70f5e05f09f93e00b997edb1ef0c864 } from '@payloadcms/richtext-lexical/client'
import { BlockquoteFeatureClient as BlockquoteFeatureClient_e70f5e05f09f93e00b997edb1ef0c864 } from '@payloadcms/richtext-lexical/client'
import { UploadFeatureClient as UploadFeatureClient_e70f5e05f09f93e00b997edb1ef0c864 } from '@payloadcms/richtext-lexical/client'
import { HorizontalRuleFeatureClient as HorizontalRuleFeatureClient_e70f5e05f09f93e00b997edb1ef0c864 } from '@payloadcms/richtext-lexical/client'
import { InlineToolbarFeatureClient as InlineToolbarFeatureClient_e70f5e05f09f93e00b997edb1ef0c864 } from '@payloadcms/richtext-lexical/client'
import { CollectionCards as CollectionCards_f9c02e79a4aed9a3924487c0cd4cafb1 } from '@payloadcms/next/rsc'
/** @type import('payload').ImportMap */
export const importMap = {
"@payloadcms/richtext-lexical/rsc#RscEntryLexicalCell": RscEntryLexicalCell_44fe37237e0ebf4470c9990d8cb7b07e,
"@payloadcms/richtext-lexical/rsc#RscEntryLexicalField": RscEntryLexicalField_44fe37237e0ebf4470c9990d8cb7b07e,
"@payloadcms/richtext-lexical/rsc#LexicalDiffComponent": LexicalDiffComponent_44fe37237e0ebf4470c9990d8cb7b07e,
"@payloadcms/richtext-lexical/client#BlocksFeatureClient": BlocksFeatureClient_e70f5e05f09f93e00b997edb1ef0c864,
"@payloadcms/richtext-lexical/client#BoldFeatureClient": BoldFeatureClient_e70f5e05f09f93e00b997edb1ef0c864,
"@payloadcms/richtext-lexical/client#ItalicFeatureClient": ItalicFeatureClient_e70f5e05f09f93e00b997edb1ef0c864,
"@payloadcms/richtext-lexical/client#UnderlineFeatureClient": UnderlineFeatureClient_e70f5e05f09f93e00b997edb1ef0c864,
"@payloadcms/richtext-lexical/client#StrikethroughFeatureClient": StrikethroughFeatureClient_e70f5e05f09f93e00b997edb1ef0c864,
"@payloadcms/richtext-lexical/client#SubscriptFeatureClient": SubscriptFeatureClient_e70f5e05f09f93e00b997edb1ef0c864,
"@payloadcms/richtext-lexical/client#SuperscriptFeatureClient": SuperscriptFeatureClient_e70f5e05f09f93e00b997edb1ef0c864,
"@payloadcms/richtext-lexical/client#InlineCodeFeatureClient": InlineCodeFeatureClient_e70f5e05f09f93e00b997edb1ef0c864,
"@payloadcms/richtext-lexical/client#ParagraphFeatureClient": ParagraphFeatureClient_e70f5e05f09f93e00b997edb1ef0c864,
"@payloadcms/richtext-lexical/client#HeadingFeatureClient": HeadingFeatureClient_e70f5e05f09f93e00b997edb1ef0c864,
"@payloadcms/richtext-lexical/client#AlignFeatureClient": AlignFeatureClient_e70f5e05f09f93e00b997edb1ef0c864,
"@payloadcms/richtext-lexical/client#IndentFeatureClient": IndentFeatureClient_e70f5e05f09f93e00b997edb1ef0c864,
"@payloadcms/richtext-lexical/client#UnorderedListFeatureClient": UnorderedListFeatureClient_e70f5e05f09f93e00b997edb1ef0c864,
"@payloadcms/richtext-lexical/client#OrderedListFeatureClient": OrderedListFeatureClient_e70f5e05f09f93e00b997edb1ef0c864,
"@payloadcms/richtext-lexical/client#ChecklistFeatureClient": ChecklistFeatureClient_e70f5e05f09f93e00b997edb1ef0c864,
"@payloadcms/richtext-lexical/client#LinkFeatureClient": LinkFeatureClient_e70f5e05f09f93e00b997edb1ef0c864,
"@payloadcms/richtext-lexical/client#RelationshipFeatureClient": RelationshipFeatureClient_e70f5e05f09f93e00b997edb1ef0c864,
"@payloadcms/richtext-lexical/client#BlockquoteFeatureClient": BlockquoteFeatureClient_e70f5e05f09f93e00b997edb1ef0c864,
"@payloadcms/richtext-lexical/client#UploadFeatureClient": UploadFeatureClient_e70f5e05f09f93e00b997edb1ef0c864,
"@payloadcms/richtext-lexical/client#HorizontalRuleFeatureClient": HorizontalRuleFeatureClient_e70f5e05f09f93e00b997edb1ef0c864,
"@payloadcms/richtext-lexical/client#InlineToolbarFeatureClient": InlineToolbarFeatureClient_e70f5e05f09f93e00b997edb1ef0c864,
"@payloadcms/next/rsc#CollectionCards": CollectionCards_f9c02e79a4aed9a3924487c0cd4cafb1
}
@@ -0,0 +1,20 @@
/*
* Payload's REST API, mounted at /cms-api rather than /api.
* See lib/payloadRoutes.ts — /api is the FastAPI proxy's catch-all.
*/
import config from '@payload-config';
import {
REST_DELETE,
REST_GET,
REST_OPTIONS,
REST_PATCH,
REST_POST,
REST_PUT,
} from '@payloadcms/next/routes';
export const GET = REST_GET(config);
export const POST = REST_POST(config);
export const DELETE = REST_DELETE(config);
export const PATCH = REST_PATCH(config);
export const PUT = REST_PUT(config);
export const OPTIONS = REST_OPTIONS(config);
@@ -0,0 +1,4 @@
import config from '@payload-config';
import { GRAPHQL_PLAYGROUND_GET } from '@payloadcms/next/routes';
export const GET = GRAPHQL_PLAYGROUND_GET(config);
@@ -0,0 +1,5 @@
import config from '@payload-config';
import { GRAPHQL_POST, REST_OPTIONS } from '@payloadcms/next/routes';
export const POST = GRAPHQL_POST(config);
export const OPTIONS = REST_OPTIONS(config);
+27
View File
@@ -0,0 +1,27 @@
/**
* Root layout for the Payload admin panel.
*
* This is a SECOND root layout: it renders its own <html>/<body>, as does
* app/(frontend)/layout.tsx. Next permits that only while no app/layout.tsx
* exists — which is why the site's routes were moved into (frontend). Adding
* an app/layout.tsx would nest the admin panel inside the site's nav, footer
* and providers and emit nested <html>.
*/
import type { ServerFunctionClient } from 'payload';
import config from '@payload-config';
import { RootLayout, handleServerFunctions } from '@payloadcms/next/layouts';
import { importMap } from './admin/importMap.js';
import '@payloadcms/next/css';
const serverFunction: ServerFunctionClient = async function (args) {
'use server';
return handleServerFunctions({ ...args, config, importMap });
};
export default function PayloadLayout({ children }: { children: React.ReactNode }) {
return (
<RootLayout config={config} importMap={importMap} serverFunction={serverFunction}>
{children}
</RootLayout>
);
}
+7 -2
View File
@@ -12,9 +12,14 @@ export default function robots(): MetadataRoute.Robots {
{ {
userAgent: '*', userAgent: '*',
allow: '/', allow: '/',
disallow: ['/api/', '/_next/'], // /admin and /cms-api are also served X-Robots-Tag: noindex by
// next.config.mjs. A Disallow alone blocks crawling, not indexing.
disallow: ['/api/', '/_next/', '/admin/', '/cms-api/'],
}, },
], ],
sitemap: absoluteUrl('/sitemap.xml'), // Two sitemaps: /sitemap.xml is proxied from FastAPI and carries the
// school corpus; /content-sitemap.xml is Next-owned and carries /about
// and the blog. The backend knows nothing about Payload.
sitemap: [absoluteUrl('/sitemap.xml'), absoluteUrl('/content-sitemap.xml')],
}; };
} }
+25
View File
@@ -0,0 +1,25 @@
import type { Block } from 'payload';
/**
* The house block: "what this number doesn't tell you".
*
* Blocks are the reason this site runs a CMS rather than flat files — a post
* can carry live product components, not screenshots of them. This is the
* first and simplest one; a live-chart block follows when a post needs it.
*/
export const Callout: Block = {
slug: 'callout',
labels: { singular: 'Callout', plural: 'Callouts' },
fields: [
{
name: 'tone',
type: 'select',
defaultValue: 'caveat',
options: [
{ label: 'Caveat: what this does not show', value: 'caveat' },
{ label: 'Note: useful aside', value: 'note' },
],
},
{ name: 'body', type: 'textarea', required: true },
],
};
+31
View File
@@ -0,0 +1,31 @@
import type { CollectionConfig } from 'payload';
/**
* Uploads land on a Docker named volume mounted at /app/media. The path is
* absolute because Payload 3 requires it, and it must match the payload_media
* mount in docker-compose.portainer.yml exactly — a mismatch writes into the
* container's own filesystem, where the next redeploy silently discards it.
*/
export const Media: CollectionConfig = {
slug: 'media',
access: { read: () => true },
upload: {
staticDir: '/app/media',
mimeTypes: ['image/*'],
imageSizes: [
{ name: 'thumbnail', width: 400 },
{ name: 'hero', width: 1200 },
],
adminThumbnail: 'thumbnail',
},
fields: [
{
name: 'alt',
type: 'text',
required: true,
// Required rather than optional: a decorative-by-default image is an
// accessibility regression on a site parents use under time pressure.
admin: { description: 'Describe the image for screen readers.' },
},
],
};
+99
View File
@@ -0,0 +1,99 @@
import type { CollectionConfig } from 'payload';
import { revalidatePath } from 'next/cache';
import { lexicalEditor, BlocksFeature } from '@payloadcms/richtext-lexical';
import { Callout } from '@/blocks/Callout';
/**
* Drop the cached copy of a post page when it changes.
*
* Only the post page needs this. The blog index, the RSS feed and the content
* sitemap are force-dynamic — they have no dynamic params, so Next would
* prerender them at build time, where CI has no database — which means they
* already reflect a change on the next request.
*
* /blog/[slug] is ISR: generated on first request and cached, so without this
* an edit to an already-published post would not appear until the revalidate
* window expired — up to an hour of a writer concluding that saving is broken.
*
* Payload runs in the same process as Next, so this is a direct revalidatePath
* call: no webhook, no shared secret, no network hop to get wrong.
*/
function revalidatePost(slug: string) {
revalidatePath(`/blog/${slug}`);
}
export const Posts: CollectionConfig = {
slug: 'posts',
access: {
/*
* Drafts must be hidden here, not in the pages that query this collection.
*
* From Payload's own documentation: "The `draft` argument alone does not
* restrict documents with `_status: 'draft'` from being returned by the
* API." The blog index and post page both filter on `_status`, but that
* is a convenience, not a control — a direct GET /cms-api/posts would
* hand every unpublished draft to any visitor.
*
* Returning a query constraint rather than a boolean is the documented
* mechanism: Payload merges it into every read for an anonymous caller.
*/
read: ({ req }) => {
if (req.user) return true;
return { _status: { equals: 'published' } };
},
},
admin: {
useAsTitle: 'title',
defaultColumns: ['title', 'publishedAt', '_status'],
},
versions: {
// Posts get written across several sittings and previewed before they go
// live. Without drafts, saving is publishing.
drafts: true,
},
hooks: {
afterChange: [({ doc }) => { revalidatePost(String(doc.slug)); }],
afterDelete: [({ doc }) => { revalidatePost(String(doc.slug)); }],
},
fields: [
{ name: 'title', type: 'text', required: true },
{
name: 'slug',
type: 'text',
required: true,
unique: true,
index: true,
admin: {
position: 'sidebar',
description: 'The URL segment. Never change it after publishing.',
},
},
{
name: 'publishedAt',
type: 'date',
required: true,
admin: { position: 'sidebar', date: { pickerAppearance: 'dayOnly' } },
},
{
name: 'excerpt',
type: 'textarea',
required: true,
maxLength: 200,
admin: {
description: 'Shown on the index and used as the meta description.',
},
},
{ name: 'heroImage', type: 'upload', relationTo: 'media' },
{
name: 'content',
type: 'richText',
required: true,
editor: lexicalEditor({
features: ({ defaultFeatures }) => [
...defaultFeatures,
BlocksFeature({ blocks: [Callout] }),
],
}),
},
],
};
+33
View File
@@ -0,0 +1,33 @@
import type { CollectionConfig } from 'payload';
/**
* The site's only authenticated surface. There is one account and no
* registration: `create` is closed to everyone, so the first user is seeded
* with `payload create-first-user` and no one can add another through the API.
*/
export const Users: CollectionConfig = {
slug: 'users',
auth: {
// Slows credential stuffing against a panel that is on the public
// internet. Five attempts, then a ten-minute lock.
maxLoginAttempts: 5,
lockTime: 10 * 60 * 1000,
},
access: {
create: () => false,
read: ({ req }) => Boolean(req.user),
update: ({ req }) => Boolean(req.user),
delete: () => false,
},
admin: { useAsTitle: 'email' },
fields: [
{
name: 'displayName',
type: 'text',
required: true,
// Rendered as the byline on every post. First name only — the site
// publishes no surname and no employer.
defaultValue: 'Tudor',
},
],
};
+21 -1
View File
@@ -48,6 +48,8 @@
display: flex; display: flex;
align-items: center; align-items: center;
gap: 0.5rem; gap: 0.5rem;
/* The suggestion dropdown is absolutely positioned against this box. */
position: relative;
} }
/* The hero pill: hairline, soft corner, everything else sits inside it. */ /* The hero pill: hairline, soft corner, everything else sits inside it. */
@@ -411,7 +413,17 @@
/* ── Narrow ───────────────────────────────────────────────────────── */ /* ── Narrow ───────────────────────────────────────────────────────── */
@media (max-width: 768px) { @media (max-width: 768px) {
.filterBar { /*
* Scoped, like the two rules below it.
*
* The results filter bar is a card — background, border, shadow — and needs
* inner padding. The hero's search is not a card: .heroMode zeroes the
* padding, border and background so the search sits directly on the panel.
* Unscoped, this rule put 14px back, which indented the search box, the hint
* and the location link 14px past the headline they sit under, and cost the
* search field 28px of width on a 390px screen.
*/
.filterBar:not(.heroMode) {
padding: 0.875rem; padding: 0.875rem;
} }
@@ -455,6 +467,14 @@
align-items: flex-start; align-items: flex-start;
} }
/* Optical alignment: the button's own 6px of padding is what makes its
label start further right than the hint above it, even once both boxes
share a left edge. Pulling the padding back off lines the text up while
keeping the tap target. */
.heroMode .nearMeBtn {
margin-left: -0.375rem;
}
.geoError { .geoError {
text-align: left; text-align: left;
} }
+84 -2
View File
@@ -3,8 +3,11 @@
import { useState, useCallback, useTransition, useRef, useEffect } from "react"; import { useState, useCallback, useTransition, useRef, useEffect } from "react";
import type { ReactNode } from "react"; import type { ReactNode } from "react";
import { useRouter, useSearchParams, usePathname } from "next/navigation"; import { useRouter, useSearchParams, usePathname } from "next/navigation";
import { isValidPostcode } from "@/lib/utils"; import { isValidPostcode, schoolUrl } from "@/lib/utils";
import { track } from "@/lib/analytics"; import { track } from "@/lib/analytics";
import { useSchoolSuggest } from "@/hooks/useSchoolSuggest";
import { SuggestList, suggestOptionId } from "./SuggestList";
import type { Suggestion } from "@/lib/suggest";
import type { Filters, ResultFilters } from "@/lib/types"; import type { Filters, ResultFilters } from "@/lib/types";
import styles from "./FilterBar.module.css"; import styles from "./FilterBar.module.css";
@@ -17,6 +20,8 @@ interface FilterBarProps {
onNearMe?: () => void; onNearMe?: () => void;
geoState?: "idle" | "requesting" | "error"; geoState?: "idle" | "requesting" | "error";
geoError?: string | null; geoError?: string | null;
/** Server-read feature flag. Off means no listener, no fetch, no markup. */
autosuggest?: boolean;
} }
/** /**
@@ -48,6 +53,7 @@ export function FilterBar({
onNearMe, onNearMe,
geoState = "idle", geoState = "idle",
geoError, geoError,
autosuggest = false,
}: FilterBarProps) { }: FilterBarProps) {
const router = useRouter(); const router = useRouter();
const pathname = usePathname(); const pathname = usePathname();
@@ -62,6 +68,59 @@ export function FilterBar({
const [omniValue, setOmniValue] = useState(initialOmniValue); const [omniValue, setOmniValue] = useState(initialOmniValue);
const suggestId = `school-suggest-${isHero ? "hero" : "bar"}`;
/*
* Suggestions answer typing, not the mere presence of a value.
*
* Without this the results-page bar reopened the dropdown over the results:
* after a search the input still holds the term, so on every render the
* query was >= 2 characters and the list opened again — on top of the very
* results the search had just produced, swallowing the click on the first
* one. The E2E gate caught it as "<li role=option> intercepts pointer
* events", but a reader would just have found the page unclickable.
*/
const [hasTyped, setHasTyped] = useState(false);
// Suppressed once the value parses as a postcode: the box takes a school
// name OR a postcode, and suggesting schools during postcode entry fights
// the user rather than helping them.
const suggestEnabled = autosuggest && hasTyped && !isValidPostcode(omniValue);
const { suggestions, open, activeIndex, setActiveIndex, close } =
useSchoolSuggest(omniValue, suggestEnabled);
const pickSuggestion = (s: Suggestion) => {
setHasTyped(false);
close();
track('search_submitted', {
query: s.school_name.toLowerCase(),
via: 'suggestion',
urn: s.urn,
has_postcode: false,
filters_active: '',
filters_count: 0,
});
router.push(schoolUrl(s.urn, s.school_name));
};
const handleOmniKeyDown = (e: React.KeyboardEvent<HTMLInputElement>) => {
if (!open) return;
if (e.key === "ArrowDown") {
e.preventDefault();
setActiveIndex(activeIndex + 1 >= suggestions.length ? 0 : activeIndex + 1);
} else if (e.key === "ArrowUp") {
e.preventDefault();
setActiveIndex(activeIndex <= 0 ? suggestions.length - 1 : activeIndex - 1);
} else if (e.key === "Escape") {
close();
} else if (e.key === "Enter" && activeIndex >= 0) {
// Only when an option is active. With none, the event falls through to
// the form's submit handler and searches the typed text, as it does now.
e.preventDefault();
pickSuggestion(suggestions[activeIndex]);
}
};
const currentLA = searchParams.get("local_authority") || ""; const currentLA = searchParams.get("local_authority") || "";
const currentType = searchParams.get("school_type") || ""; const currentType = searchParams.get("school_type") || "";
const currentPhase = searchParams.get("phase") || ""; const currentPhase = searchParams.get("phase") || "";
@@ -124,6 +183,9 @@ export function FilterBar({
const handleSearchSubmit = (e: React.FormEvent) => { const handleSearchSubmit = (e: React.FormEvent) => {
e.preventDefault(); e.preventDefault();
// The search has been made; the suggestions that led to it are spent.
setHasTyped(false);
close();
if (!omniValue.trim()) { if (!omniValue.trim()) {
updateURL({ search: "", postcode: "", radius: "" }); updateURL({ search: "", postcode: "", radius: "" });
return; return;
@@ -226,9 +288,20 @@ export function FilterBar({
ref={inputRef} ref={inputRef}
type="search" type="search"
value={omniValue} value={omniValue}
onChange={(e) => setOmniValue(e.target.value)} onChange={(e) => { setOmniValue(e.target.value); setHasTyped(true); }}
onKeyDown={handleOmniKeyDown}
onBlur={close}
placeholder="School name or postcode" placeholder="School name or postcode"
className={styles.omniInput} className={styles.omniInput}
{...(autosuggest ? {
role: "combobox",
"aria-expanded": open,
"aria-controls": suggestId,
"aria-autocomplete": "list" as const,
"aria-activedescendant":
activeIndex >= 0 ? suggestOptionId(suggestId, activeIndex) : undefined,
autoComplete: "off",
} : {})}
/> />
<button <button
type="submit" type="submit"
@@ -237,6 +310,15 @@ export function FilterBar({
> >
{isPending ? <div className={styles.spinner}></div> : isHero ? "Search schools" : "Search"} {isPending ? <div className={styles.spinner}></div> : isHero ? "Search schools" : "Search"}
</button> </button>
{autosuggest && open && (
<SuggestList
id={suggestId}
suggestions={suggestions}
activeIndex={activeIndex}
onPick={pickSuggestion}
onHover={setActiveIndex}
/>
)}
</div> </div>
{isHero && ( {isHero && (
<> <>
+10 -1
View File
@@ -22,7 +22,8 @@
.content { .content {
display: grid; display: grid;
grid-template-columns: 1.6fr 1fr 1fr; /* Brand column plus three link columns: Product, Resources, About. */
grid-template-columns: 1.6fr 1fr 1fr 1fr;
gap: 2rem; gap: 2rem;
margin-bottom: 3rem; margin-bottom: 3rem;
} }
@@ -193,6 +194,14 @@
color: var(--on-sunken); color: var(--on-sunken);
} }
/* Four columns crush between the tablet range and the 768px collapse, so
pair them up first rather than jumping straight to a single column. */
@media (max-width: 960px) {
.content {
grid-template-columns: 1fr 1fr;
}
}
@media (max-width: 768px) { @media (max-width: 768px) {
.container { .container {
padding: 2rem 1rem 1.5rem; padding: 2rem 1rem 1.5rem;
+32 -1
View File
@@ -10,7 +10,17 @@
import { LogoMark } from './Logo'; import { LogoMark } from './Logo';
import styles from './Footer.module.css'; import styles from './Footer.module.css';
export function Footer() { /**
* Both default to false so a caller that forgets a prop hides the link rather
* than pointing it at a page that 404s. Same reasoning as backend/flags.py:
* "Every flag defaults to False."
*/
interface FooterProps {
aboutEnabled?: boolean;
blogEnabled?: boolean;
}
export function Footer({ aboutEnabled = false, blogEnabled = false }: FooterProps = {}) {
const currentYear = new Date().getFullYear(); const currentYear = new Date().getFullYear();
return ( return (
@@ -93,6 +103,27 @@ export function Footer() {
</li> </li>
</ul> </ul>
</div> </div>
{/* Dropped entirely when both flags are dark, rather than left as an
empty heading: shipping dark means the footer renders as it did
before the feature existed. */}
{(aboutEnabled || blogEnabled) && (
<div className={styles.section}>
<h4 className={styles.sectionTitle}>About</h4>
<ul className={styles.links}>
{/* The only route to a named human. Deliberately not in the nav:
the mobile bottom bar already carries four items, and both of
these are lower intent than any of them. Post bylines link
here too, which is where a reader actually asks the question. */}
{aboutEnabled && (
<li><a href="/about" className={styles.link}>Who&apos;s behind this</a></li>
)}
{blogEnabled && (
<li><a href="/blog" className={styles.link}>Blog</a></li>
)}
</ul>
</div>
)}
</div> </div>
<div className={styles.bottom}> <div className={styles.bottom}>
+23 -1
View File
@@ -90,7 +90,18 @@
isolation: isolate; isolation: isolate;
background: var(--hero-ground); background: var(--hero-ground);
border-radius: var(--radius-xl); border-radius: var(--radius-xl);
overflow: hidden; /*
* Deliberately NOT overflow: hidden.
*
* It used to be, to clip the artwork and the scrim to the rounded corners —
* and it also clipped the search box's suggestion dropdown, which is 320px
* tall against 145px of panel below the input. Roughly half the list was cut
* off with no indication anything was missing.
*
* The two things that actually needed clipping round themselves instead, so
* the panel can let a dropdown out. Anything absolutely positioned inside
* this panel and taller than the space below it depends on this.
*/
} }
.heroContent { .heroContent {
@@ -107,6 +118,10 @@
position: absolute; position: absolute;
inset: 0; inset: 0;
z-index: 0; z-index: 0;
/* Rounds itself, because the panel no longer clips it. inset: 0 makes this
exactly the panel's own corners. */
border-radius: inherit;
overflow: hidden;
} }
.heroArt picture, .heroArt picture,
@@ -143,6 +158,9 @@
inset: 0; inset: 0;
z-index: 1; z-index: 1;
pointer-events: none; pointer-events: none;
/* Same reason as .heroArt: the panel stopped clipping, so the scrim keeps
its own corners rather than squaring off over the panel's. */
border-radius: inherit;
background: linear-gradient( background: linear-gradient(
to right, to right,
var(--hero-ground) 0%, var(--hero-ground) 0%,
@@ -331,6 +349,10 @@
position: static; position: static;
order: -1; order: -1;
height: 13rem; height: 13rem;
/* Top corners only. Here the artwork is a band flush with the top of the
panel, not a layer covering it — inheriting all four would leave it
floating with rounded bottom corners against the copy below. */
border-radius: var(--radius-xl) var(--radius-xl) 0 0;
} }
/* The band crop puts the schoolhouse at 73% across — reported by /* The band crop puts the schoolhouse at 73% across — reported by
scripts/build-hero-images.js, which derives it from the crop box rather scripts/build-hero-images.js, which derives it from the crop box rather
Loaded 100 of 171 files, more files were not shown because too many files have changed in this diff. Show more