The site reads as synthetic because nobody is accountable for the
numbers, no editorial judgement is visible, and the voice is
institutional third person. This designs the fix: a named author
(first name, photo, explicitly not an education expert), a coded
/about page, and a blog backed by Payload CMS running inside the
existing Next app.
Also fills a hole in the SEO programme, which has eight workstreams
and no E-E-A-T or authorship signal on a YMYL corpus.
Records two collisions found while designing, both of which fail
badly if missed: Payload's default /api route fights the existing
FastAPI catch-all proxy, and withPayload() is ESM-only so
next.config.js — which carries staging's noindex header — has to
become next.config.mjs.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FT1Ls4GbgLDXoQX7NAuHGT
Review found _mask_for_disclosure could return with its invariant broken
and say nothing. add_companion only ever withheld a *published* cell, so a
group with one suppressed category and every other one not_applicable —
routine in special schools and AP, where few categories apply — left the
loop with the lone suppressed cell still solvable. Reproduced on a
nine-pupil cohort: one hidden cell, cohort served, residual intact.
A disclosure-control pass that fails silently is worse than none, because
everything downstream trusts it. The loop now runs until the invariant
holds and escalates when no companion exists: the pupil group is dropped
from the payload, and an empty block serialises as None so the section is
absent rather than an empty shell. disclosure_invariant_holds() is exported
so tests assert it directly instead of re-deriving it, and an exhaustive
test sweeps all 81 suppression patterns of a four-category group.
Also fixes a test that set up six measures and checked one: the loop was
`for measure in ["school_sixth_form"]`. It now checks every measure, and
against the real invariant — none hidden, or at least two, rather than
"at least two", which the five published measures would have failed.
No regression on real data: 262 mainstream secondaries, all-pupils bar
still drawable on 94%, zero invariant violations, one disadvantaged group
dropped by the new escalation.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvdDKvFFSZuMVDH5fEyTob
Code review found the disclosure the whole design was meant to prevent.
R1 was written as a rendering rule and implemented as one: canRenderBar
stopped the bar being drawn, but GET /api/schools/{urn} still carried the
cohort and every published category. cohort - sum(published) returned
Whitley Bay's withheld further-education figure exactly — 18 pupils — to
any caller, and the RSC payload put it in the browser too.
app.py already stated the principle for admission_distance: this endpoint
is public and unauthenticated, so a field left in the payload is a
published field. The same reasoning applies here and did not get applied.
_mask_for_disclosure now closes both identities before serialisation —
categories sum to the cohort, and disadvantaged + other = all — by adding
secondary suppression until every row and column hides none or at least
two. My first attempt picked the smallest published cell as the companion
and a new test caught it choosing a zero, which protects nothing: the
residual still resolved to 18. The companion must carry pupils.
DfE's own aggregates are no longer served. Nothing rendered them, and one
spanning a single suppressed component names it.
Cost, measured over 262 mainstream secondaries: the all-pupils bar
survives on 94% rather than 100%. Zero lone-suppressed groups remain.
The e2e helper now tells a missing feature apart from missing data: it
fails if the API serves no destinations key at all, and skips if the key
is served but the annual DAG has not populated the marts. Failing on the
second would redden the staging gate for unrelated commits.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvdDKvFFSZuMVDH5fEyTob
The published files suppress individual cells, not whole cohorts, and the
categories sum to the cohort — so on 22% of mainstream secondaries the
withheld figure can be recovered by subtraction. Three disclosure rules
fall out of that, and the rest of the design is downstream of them.
Verified against the EES API rather than assumed: both datasets carry
school-level rows keyed by URN, with the disadvantage split and every
destination category the display needs.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvdDKvFFSZuMVDH5fEyTob
Code review, both findings valid.
The design doc claimed Cloudflare "replaces the header, so a browser
cannot forge it", and that only the X-Forwarded-For fallback was
forgeable. That is true only for traffic that actually passed through
Cloudflare, and nothing in this process can verify that it did. Reaching
the origin directly, both headers are equally attacker-controlled — and
rotating CF-Connecting-IP mints a fresh rate-limit bucket per request,
defeating per-client limits on every endpoint including the
DataFrame-heavy /api/schools. Against abuse that is worse than the
shared bucket it replaced, which at least capped everyone together.
So the ceiling comes back. I dropped it earlier arguing it belonged at
Cloudflare; that argument assumed the keying was sound, and it is not.
GlobalRateLimitMiddleware counts all /api/ traffic in a fixed window
against a total, independent of client identity, outermost so it refuses
before any work happens. Written by hand because slowapi cannot express
a global cap: default_limits and application_limits are both keyed by
key_func, and the latter needs middleware this app does not install.
It does not make the header trustworthy — it makes trusting it
survivable. The real fix is Authenticated Origin Pulls or an origin
firewall, now documented in DEPLOY.md as the open gap it is.
127.0.0.1 is exempt: the healthcheck curls localhost from inside the
container, and starving it would restart the container and turn a load
spike into an outage loop. Keyed on the peer address, never the Host
header, which the caller sets.
Second finding: suggest_schools_typesense promised "never raises" while
the parsing loop sat outside the try, so int(None) on a malformed
document would have made a keystroke a 500. The loop now skips bad rows
rather than dropping the whole list — and a hit with no document no
longer becomes a suggestion pointing at /school/0.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
Reading slowapi rather than assuming: default_limits and
application_limits are both evaluated with the same key_func, so they
are per-client across routes, not global. And application_limits only
apply 'if in_middleware' — this app installs no SlowAPIMiddleware, so
they would never have fired at all.
A genuine global cap would need a second Limiter with a constant key
plus that middleware. Cloudflare is already in the path on both
environments and does this at the right layer, so the ceiling is named
as a follow-up there rather than built badly here.
The risk that leaves is stated plainly in the risks section instead of
being papered over with a mechanism that does not do the job.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
The load-bearing finding is not about autosuggest. The rate limiter keys
on request.client.host, which in staging and prod is the Next container
— so all browser users share one 60/min bucket per route. Measured
against staging: 70 concurrent requests gave exactly 60 x 200 and
10 x 429. Eight concurrent searchers would 429 the site once each
keystroke costs a request, so the keying fix is part of this work.
Both environments are behind Cloudflare, which sets CF-Connecting-IP and
overwrites any client-supplied value — trustworthy in a way a parsed
X-Forwarded-For chain is not, and the backend is unreachable except
through the Next proxy.
Named honestly: the shared bucket has been an accidental global throttle
on a single-process backend, so correct per-user keying removes a
protection. A global ceiling ships with it rather than instead of it.
Suggestions come from Typesense alone. The existing search path filters
a 25,000-row DataFrame per query, which is exactly the cost a keystroke
endpoint cannot pay, so there is deliberately no DataFrame fallback.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
Next uses the LOWEST revalidate among a route's fetches, not the segment
value. School pages fetch school details at 300s and place pages fetch
national averages at 3600s, so the effective ISR period is five minutes
and one hour respectively — not the seven days the segment declares.
A flag flip therefore propagates on its own, well inside the monthly,
by-hand cadence these flags are for. That deletes two webhook
integrations, a revalidate route, a secret-in-query-string scheme, an
idempotency requirement, and the rule that every fetch carry a cache
tag — which was the part most likely to rot as fetches are added.
Two constraints survive: a flag must never gate content on a
force-static page, because app/admissions never revalidates; and a
route-family flag must rebuild the sitemap, deferred with the route case
since no flag in scope touches it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
Unleash self-hosted in its own Portainer stack, with FastAPI holding the
only SDK and Next reading flags through a tagged fetch.
The two hard parts are consequences of putting flag state in a service
rather than the repo: main stops being the whole truth about what is on,
and a flag can now change without the deploy that would have cleared the
caches. A code-declared registry bounds the first; webhook-driven
revalidateTag handles the second.
Cache tagging is deliberately coarse — every server fetch carries the
flags tag, not just the flags fetch itself. The first consumer proves
why: admission_distance changes the shape of /api/schools/{urn}, so a
narrow purge would leave ~25,000 school pages serving the pre-flip
render for a week, invisibly.
First consumer is the last-distance-offered feature, which is on main
and staging and has never reached production. It needs one gate, at the
API, because the frontend already no-ops on a missing field.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
Supersedes the original spec's W2. The Search Console baseline inverted its
ordering: every measured location query is town or district level, none is an
administrative area, and phase is part of the query rather than a filter.
Two problems the original design did not anticipate. 67 viable towns share a
name with a local authority, and the authority is the larger set in only 43 of
them — postal towns cross authority boundaries, so neither can absorb the
other. Two namespaces resolve it by construction. And the GIAS town field
collapses 1,819 London schools into one value, which a curated
locality-to-outcode seed solves without new ingestion.
Sizing is measured against the live 25,185-school corpus rather than
estimated: 783 viable towns, 1,760 outcodes, 154 authorities.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
GIAS ships the whole UK plus overseas and offshore establishments. None of
them carry comparable DfE performance data — Wales does not publish on the
English measures at all — so every one of these pages rendered with null
results, null Ofsted and null phase. There were 2,036 of them: 1,569 Welsh,
123 offshore (Jersey, Guernsey, Isle of Man, Gibraltar), 316 British schools
overseas and 28 service children's schools. All 2,036 were being submitted to
search engines, alongside 29 local authorities that existed in the filters
purely to list them.
Filter at the mart boundary rather than the view layer. dim_school and
dim_location both exclude TypeOfEstablishment in {25, 26, 30, 37}, listed once
as vars.non_england_school_type_codes. Everything downstream reads those two
marts — search, the school page, /api/filters, rankings, Typesense and
build_sitemap() — so one filter removes them from the site and the sitemap
together, and Typesense drops them on its next rebuild since it recreates the
collection and swaps the alias rather than upserting in place.
coalesce rather than a bare NOT IN: a null type code would make the predicate
null and drop the row silently, and an unknown type is not grounds for
exclusion. No establishment has a null type today, but a future GIAS refresh
could ship one and the loss would be invisible.
assert_england_only_schools guards both directions: no excluded type survives
in dim_school, and dim_location holds no URN dim_school lacks — the API
inner-joins them, so the two filters drifting apart would silently shrink the
corpus.
Corpus goes from 27,229 schools to 25,193, and the authority list from 182 to
153. The 1,569 Welsh URLs now 404.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
Two unrelated working-tree changes, committed as one at the owner's request.
Removes "Built for parents, by a parent." and its comment from the hero. It
was added two commits ago; taking it out is the owner's call, and the reasons
it was placed under the search rather than in the footer no longer apply.
The .heroByline rules in HomeView.module.css are deliberately left in place.
They are now unreferenced, but the class was purpose-built for this one line
and keeping it makes restoring the byline a one-line change. If the removal is
permanent, that block (and its 640px media query) should go with it.
Also updates two quoted figures in the 2026-07-02 UX audit notes from
"24,000+" to "27,000+".
Worth noting for the record: that file documents what the page said when it was
audited, and at that time it genuinely did say "24,000+" — which was itself the
bug later fixed by reading unique_schools instead of a field the API never
sent. Editing the quoted evidence makes the note read consistently with the
current site, at the cost of no longer being a verbatim record of what was
observed. Left as the owner edited it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Nine tasks mapping to the spec's commit sequence, each ending in an
independently testable deliverable.
Also corrects the spec's section-sharing table: measured similarity shows
Admissions (14%) and History (40%) are not shareable between the two views,
only Ofsted (80%) and Finances (91%). This costs no bytes -- the win comes
from sections being server components, not from sharing them -- and cuts the
diff and regression risk.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Planning surfaced a gap: the two detail views' CSS modules share 79 class
names, 22 with differing rules, 16 of those used in section markup by both
views. Shared sections cannot use one stylesheet without visual change.
Resolution: union the 14 incidental-drift differences (defensive overflow
properties never back-ported between the views), and keep the 2 genuine
visual differences (.genderBar, .heroStatValue) as distinct classes behind a
variant prop. Adds an isolated CSS-merge commit so a visual regression
bisects to the merge rather than a JSX move.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Design for splitting SchoolDetailView (65KB) and SecondarySchoolDetailView
(46KB) into server section components behind a small client shell, following
the existing components/compare/ decomposition pattern.
Key constraint: a server component imported by a client component becomes
client, so sections are composed in page.tsx and passed through the shell as
children. Moving /api/national-averages server-side is a prerequisite, since
the England-comparison deltas feed most sections.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The final review claimed DfE published school-level 2021/22 KS2 in Dec
2022. Re-verified: the GOV.UK announcement 'Primary school performance
tables: 2022' is CANCELLED ('will not be published in key stage 2
performance tables in academic year 2021/22'), and the EES 2021/22
release carries the same statement. No code change.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0146VHeLAWjDVE2B5uU67jCB
Recheck with 'B91 3DL' via the search box shows the geocoded radius
search works (distances, nearest-first, radius selector, map toggle).
Original P0.3 tested only a partial postcode, which silently falls back
to text search — refiled as P2.0 (silent mode switch on partial input).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>