Compare commits

...
Author SHA1 Message Date
TudorandClaude Opus 5 1d8858fbda chore: remove the code the legacy CSV importer left behind
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m11s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 18s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m18s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 11s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 1m2s
`backend/migration.py` and `scripts/migrate_csv_to_db.py` import `School`,
`SchoolResult`, `init_db` and `set_db_schema_version` — names that no longer
exist. `scripts/geocode_schools.py` imports the same removed ORM model. None of
the three can be imported against the current backend, so they were not dormant
utilities anyone could fall back on; they were files that would fail on the
first line. `backend/version.py` existed only to hand `SCHEMA_VERSION` to that
importer, and the FastAPI lifespan performs no version-triggered import.

Three symbols go with them, each confirmed to have no caller: the unvectorised
`haversine_distance`, superseded by the inline NumPy calculation in search;
`fetcher`, an SWR helper for a dependency this project does not install; and
`kmToMiles`. `calculateDistance` stays — CutoffMapPanel uses it.

Two comments pointed at `migrate_csv_to_db.py --drop` to explain why Payload
owns its own schema. The reason survives the script: blog content must stay
clear of the school marts and Airflow's metadata. Reworded rather than deleted,
so the constraint keeps its justification.

docs/LEGACY_CODE.md records what was removed and where to find it in history. It
also records what was deliberately *not* removed, which is the more useful half:
unused UI components awaiting a design decision, manual data utilities whose
operators a repository search cannot see, and fallbacks that look obsolete but
are load-bearing — `data_loader.py`'s older-mart branches, the generated GIAS
dictionary copies, and the `legacy`-named dbt models that annual DAG selectors
explicitly include. A zero-import count is evidence, not a verdict.

The scripts that fetch DfE CSVs are marked historical and kept, pending
confirmation that nobody runs them by hand.

Checked: 190 backend tests, 429 frontend tests, `tsc --noEmit` clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016y2J6bs8gbuSJbH18w7Tan
2026-09-14 23:01:22 +01:00
TudorandClaude Opus 5 eaf5e5d180 docs: describe the system that exists, not the one we started with
The README still opened on "Primary School Compass", a KS2 tool for Wandsworth
and Merton served by FastAPI and vanilla JavaScript with Chart.js. Every layer
of that sentence is now wrong: coverage is England-wide across KS2, KS4,
all-through and post-16, Next.js owns the public UI, and school data comes from
dbt-built `marts.*` rather than CSVs loaded at startup. The setup instructions
walked a reader into a virtualenv and a CSV import that cannot build the current
schema, so following the docs produced an empty database and a wrong mental
model at the same time.

Replace the narrative docs with two reference documents that were checked
against the code: docs/ARCHITECTURE.md for request flow, data ownership, the
backend/frontend module boundaries and the real publication sequence, and
docs/DEVELOPMENT.md for the checks that actually run, including the container
and CI version skew that makes "just run pytest" misleading.

The env examples drifted the same way. ALLOWED_ORIGINS is a JSON array, not a
comma-separated list; the frontend needs FASTAPI_URL, DATABASE_URL and
PAYLOAD_SECRET, none of which were documented; and RATE_LIMIT_BURST,
DEFAULT_PAGE_SIZE and MAX_PAGE_SIZE were presented as tuning controls the routes
do not consult. Each is now stated as it behaves.

MIGRATION_SUMMARY.md keeps its content but gains a banner, because it reads like
setup instructions and is not.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016y2J6bs8gbuSJbH18w7Tan
2026-09-14 23:01:15 +01:00
tudor be780ebe13 Merge pull request 'fix(seo): declare the share card, which the route group stopped inheriting' (#146) from fix/og-image-route-group into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 1m21s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Successful in 1m58s
Reviewed-on: #146
2026-09-14 21:03:14 +00:00
TudorandClaude Opus 5 b6c2cd5116 fix(seo): declare the share card, which the route group stopped inheriting
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m11s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 11s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m17s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 11s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 53s
The staging E2E gate's one failure. Every link to the site pasted into
a chat has been rendering bare.

Staging serves og:title, og:description, og:url, og:site_name, og:type
and twitter:card, and no og:image at all. So metadata from the layout
reaches the page; only the file convention does not.

app/opengraph-image.tsx does work — _not-found, which lives in the app
root segment, carries an og:image from it in the build output. It does
not reach the site's pages, which live in the (frontend) route group
whose own layout.tsx is the root layout. The icon conventions are not
affected: /icon.png and /apple-icon.png are both linked correctly on
the same page, verified against staging. The asymmetry is the whole
bug, and it arrived with the route-group split that Payload required.

The file stays at the app root. Moving metadata files into a route
group is what drops /robots.txt and hashes /icon.png, which CLAUDE.md
records and which this must not undo — the build still emits all four
of /robots.txt, /icon.png, /apple-icon.png and /opengraph-image. The
root layout points at the route instead, and metadataBase makes it
absolute, which the journey needs since it calls new URL() on the value.

twitter.images is set for the same reason: the card is declared
summary_large_image, and claiming a large-image card while supplying no
image is worse than claiming a summary card.

Checked before fixing that og:image was the only broken assertion in
that journey: the test aborts at line 911, so its apple-touch-icon and
maskable-icon assertions had never run. All four of those assets return
200 image/png from staging, so this does not simply move the failure
further down the test.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DXnXQKnPpZBBP61fBQiFkq
2026-09-14 21:54:32 +01:00
tudor dc79d653e5 Merge pull request 'feat(seo): link school pages into the location layer' (#145) from feat/school-page-place-links into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 21s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 1m26s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 2m29s
Reviewed-on: #145
2026-09-14 20:29:48 +00:00
TudorandClaude Opus 5 b0d5334e06 perf(places): index the reverse lookup, and isolate the registry in tests
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m11s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 18s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m17s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 11s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 1m47s
Two review findings, both confirmed before fixing.

The client fixture in test_school_details.py patched load_school_data
but not _place_registry, which is a module-level cache. A probe settled
it rather than an argument: poisoning the global with a registry built
from data the fixture never saw, then issuing the fixture's own
request, returned that other dataset's places. So the new
`places == []` assertion was satisfied by a stale registry exactly as
well as by the fixture's own data, and proved nothing. Every other test
that touches place data already reset it; the fixture predates places
existing and was never updated. It resets it now.

places_for_urn walked every place in the registry and did a tuple
membership test against each, on /api/schools/{urn}, the site's
highest-traffic endpoint. It now reads a dict built once per registry.
Measured against a synthetic corpus of 27k schools in 1,650 places:
0.118ms per request becomes 0.0001ms, with the index built once in
21ms. Production carries ~5,000 places, so the scan there is larger
again. The absolute saving per request is small; the point is that it
is repeated on every school page view and costs nothing to remove.

The index is cached against the registry by identity rather than
behind a second flag. Anything that drops _place_registry — every test
that touches place data does — gets a fresh registry object, which no
longer matches what the index was built from, so the index rebuilds
with it. A separate _place_index = None would be one more thing to
forget, and a stale reverse index is precisely the first finding's bug
wearing a different hat.

That invalidation has its own test, and the test was checked by
breaking the identity check: five tests fail without it, so three
existing ones were already relying on it too.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DXnXQKnPpZBBP61fBQiFkq
2026-09-14 21:22:39 +01:00
TudorandClaude Opus 5 d65eb58883 fix(seo): build the phase links the docstring already promised
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m12s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 18s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m18s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 11s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 3m48s
Review caught _places_payload documenting a `phase_url` the function
never returned. The docstring was not stray prose: the approved design
included the phase variant — "Primary schools in Beccles" was one of
its four example links — and it was dropped during implementation
without being mentioned. Deleting the sentence would have closed the
report while losing the feature, so the links are built instead.

These are the pages that most needed them. ~950 phase variants were
once reachable by nothing at all: absent from every sitemap and
unlinked from the place page. "Primary schools in brentwood" is the
query they exist to answer.

Membership is read from the registry's own `phase_urns` rather than
re-derived from the school's phase string. The registry already decides
which phases a place publishes and which schools are listed on each, so
asking it is both shorter and the only way the link cannot disagree
with the page it points at. It also means outcodes need no special
case: they carry empty `phase_urns` by design, because nobody searches
"primary schools in SW11", so they report no phase links on their own.

`phases` is a list rather than a single url. An all-through school is
listed on both the primary and secondary pages, so there is no tie to
break and no reason to invent one. Each entry renders directly after
its own place, so "22 primary schools in Brentwood" reads as part of
Brentwood rather than as an unrelated link further along the row.

The e2e journey now follows a phase link where the town publishes one
and asserts it resolves.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DXnXQKnPpZBBP61fBQiFkq
2026-09-14 20:54:59 +01:00
TudorandClaude Opus 5 7f5f0fb676 feat(seo): link school pages into the location layer
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m14s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 22s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m24s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 11s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 1m18s
W2 shipped ~5,000 place pages and nothing linked into them. The
location layer pointed down at school pages; school pages pointed
nowhere on the site. Their only anchor was the school's own website, so
the ~27k pages carrying most of the site's inbound authority passed it
straight off-site, and the new corpus was reachable mainly through the
sitemap.

Three things close the loop.

A reverse index over the place registry, places_for_urn, answers which
published places contain a school. Derived from the registry rather
than stored beside it, so the two cannot disagree about which places
exist: a place below the publish threshold is absent from the registry
and therefore never offered as a link. A test asserts that invariant
across every place in a built registry.

GET /api/schools/{urn} gains a `places` array carrying the name, count
and canonical path for each. It rides on the request the page already
makes, so the school page costs no extra round trip. The frontend types
it optional and defaults it to empty, because the two images deploy
separately and a frontend ahead of the API must render without it.

The page gains a "More schools near here" module and a BreadcrumbList.
The module orders narrowest first, because a reader on a school page
wants its town before its county, while the API orders widest first for
the trail. Anchors state their destination's size — "37 schools in
Brentwood" — which is worth more to a reader and a crawler than "see
more". With no published places it renders nothing rather than an empty
heading.

The trail is rooted at the homepage, not /schools. There is no /schools
index page; the location layer lives only at /schools/[place],
/schools/authority/[la] and /schools/near/[outcode]. Rooting it at the
bare path would have opened every breadcrumb with a link to a 404.
Outcodes are omitted from the trail: "schools near CM15" is a real
query and a useful link, but nobody navigates Essex to CM15 to a
school, and a breadcrumb claiming that describes a hierarchy the site
does not have.

School pages also now declare the School type rather than
EducationalOrganization, the parent type that covers universities and
nurseries alike.

The e2e journey asserts the round trip in both directions, following a
place page's own first school so the pair is genuinely related rather
than hardcoded. A one-way link is what already existed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DXnXQKnPpZBBP61fBQiFkq
2026-09-14 20:45:50 +01:00
tudor 47f3591ed8 Merge pull request 'feat(flags): put /about and /blog behind flags, dark by default' (#144) from feat/about-blog-flags into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 20s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 1m22s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 0s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 2m10s
Reviewed-on: #144
2026-09-08 16:47:18 +00:00
TudorandClaude Opus 5 6fc7fce948 feat(flags): put /about and /blog behind flags, dark by default
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m15s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 20s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m21s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 12s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 1m39s
Both features ship dark. Neither is reachable in an environment where
its flag is off, and every flag in this system starts off, so a deploy
of this commit makes both disappear until someone turns them on
deliberately.

Two independent flags rather than one, which makes blog-on-about-off a
reachable state. That state is the whole reason the change is larger
than four notFound() calls: the blog leans on the About page for its
author identity. The Person entity is anchored at /about#tudor, and
that URL 404s while about_page is dark, so a post published in that
state would claim an author resolving to nothing. Worse than having no
named author. Both bylines fall back to unlinked text and the
BlogPosting attributes to the publisher instead, so every combination
of the two flags renders something correct.

Gated: /about, /blog, /blog/[slug], the RSS feed, both footer links,
and the matching content-sitemap entries. A sitemap must never
advertise a URL that 404s. With both dark it emits a valid empty
urlset rather than a 404, because robots.txt names it unconditionally.

Not gated: /admin. Posts have to be writable before the blog is worth
switching on, so flagging the panel would make the flag unflippable.

getFlags takes a revalidate rather than always using the 300s
constant. Reading a flag pins the calling route to the lowest
revalidate among its fetches, and the footer links live in the root
layout, so a naive gate there would have dropped every school and
place page from a weekly cache to a 5-minute one. The layout passes
604800, the floor those routes already declare, and the build confirms
all four SSG route families still prerender. The cost is one-way
latency: pages follow a flip in minutes, footer links within a week.

The e2e journeys follow the existing paired shape from the
admission_distance flag: a lit journey and a dark one for each flag,
reading state from whether /about and /blog respond rather than from
/api/flags, which another journey asserts is not publicly reachable.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DXnXQKnPpZBBP61fBQiFkq
2026-09-08 17:12:25 +01:00
tudor eb6d918650 Merge pull request 'fix(cms): regenerate the import map so the Content field renders' (#143) from fix/payload-import-map into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 15s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 1m13s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 2m9s
Reviewed-on: #143
2026-09-02 20:12:17 +00:00
TudorandClaude Opus 5 d47ac71c47 fix(cms): regenerate the import map so the Content field renders
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m11s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 11s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m10s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 11s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 52s
Creating a post in the admin panel showed no Content editor, and saving
failed validation on the field the writer was never shown.

The admin panel does not import field components. The server hands the
client a path per field, and resolves it through the generated map at
app/(payload)/admin/importMap.js. A richText field's path is
@payloadcms/richtext-lexical/rsc#RscEntryLexicalField. The committed map
held one entry, @payloadcms/next/rsc#CollectionCards, generated before
the blog collections existed and never re-run. A path missing from the
map is not an error the panel reports: the field simply does not render,
while required is still enforced server-side on save.

next build does not regenerate the map, so the stale copy shipped in the
image and the editor was equally broken on staging and production.

Regenerated with payload generate:importmap, which adds the lexical RSC
field, cell and diff components, BlocksFeatureClient for the Callout
block, and the default toolbar features.

Two things stop it drifting again. There was no script to run, so
package.json gets generate:importmap. And a test asserts the map carries
an entry for each thing the config asks for, in the source-reading style
of the other payload suites; against the old map all five fail.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DXnXQKnPpZBBP61fBQiFkq
2026-09-02 20:44:05 +01:00
tudor 17e5371e9c Merge pull request 'fix(cms): ship the initial Payload migration (recovers two commits stranded after #140 merged)' (#141) from fix/payload-initial-migration into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 1m13s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 2m13s
Reviewed-on: #141
2026-09-02 17:52:17 +00:00
TudorandClaude Opus 5 e2c63a9905 fix(cms): ship the initial migration so a container finds its tables
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m14s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 11s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m14s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 1m33s
Staging failed on boot with 42P01, relation "payload.users" does not
exist. The schema was empty because no migration existed, and the
adapter cannot create tables itself: db-postgres/connect.js gates push
on NODE_ENV !== 'production', so it is inert in a deployed container
regardless of config.

The generated migration is schema-qualified to "payload" throughout but
does not create that schema — schemaName says where tables go, it does
not create anything. It only worked against the throwaway database used
to generate it because the schema was created there by hand, so every
real environment would have failed on the first statement. CREATE SCHEMA
IF NOT EXISTS is hand-added at the top of up(), which makes it exactly
the kind of edit a regeneration discards silently; a test asserts it is
present and ordered before the first CREATE TABLE.

payload-types.ts is now committed rather than ignored. Ignoring it meant
CI typechecked against looser types than a developer with a generated
copy, which is how a Record<string, unknown> cast passed CI and then
failed locally the moment the file appeared. The post page uses the
generated Post and Media types instead, and narrows heroImage rather
than asserting it, since the field is an id at shallow depth.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017YmbBhr8s7GusjDE12hrZM
2026-09-02 18:33:27 +01:00
TudorandClaude Opus 5 3f3c5953f6 style(copy): remove em dashes from the site's prose
The em dash is one of the clearest tells of machine-written text, which
is the exact impression this work exists to remove. Rewritten rather
than substituted: where a dash was carrying a real aside the sentence is
split or recast, not patched with a comma.

Covers the About page, the two Callout labels an editor sees in the
admin panel, and PUBLISHING.md, which defines the house style and should
follow it. The rule is now recorded in that house style and in the
spec's voice rules, so it survives this branch.

Code comments are left alone: they are not copy, and the surrounding
codebase uses the same punctuation throughout.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017YmbBhr8s7GusjDE12hrZM
2026-09-02 18:33:27 +01:00
tudor 124c6702a9 Merge pull request 'feat: a named author, an About page and a Payload blog' (#140) from feat/about-and-blog into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 45s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 1m32s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 2m8s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 2m36s
Reviewed-on: #140
2026-09-02 16:04:58 +00:00
TudorandClaude Opus 5 e25722d9ab fix(blog): hide drafts at the access layer, and back the --drop claim
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m12s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 32s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m9s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 1m15s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 2m26s
Review findings on #140.

Drafts were reachable. Posts granted unconditional public read and the
_status filter lived only in the pages that query the collection — which
is a convenience, not a control. Payload's documentation is explicit:
"The `draft` argument alone does not restrict documents with _status:
'draft' from being returned by the API." A direct GET /cms-api/posts
would have handed every unpublished draft to any visitor. Read access
now returns a query constraint for anonymous callers, which is the
documented mechanism.

The --drop claim was asserted across four files while the spec still
listed it as an open question. Now verified rather than assumed:
run_full_migration drops exactly ["school_results", "schools"] by name,
there is no drop_all() or DROP SCHEMA anywhere in backend/, the only
other drop is schema-qualified to marts, and nothing sets search_path.
The guarantee is stronger than schema isolation alone — those two table
names do not exist in Payload — so the claim stands, but it now rests on
cited code. The spec records the evidence and closes the open item.

findPost is wrapped in React's cache(): Next calls generateMetadata and
the page separately for one request, so every post view ran the same
query against Postgres twice.

The bare .lede rule was dead — .prose p scores (0,1,1) and outranks it —
so only .prose .lede ever applied. Removed, with the specificity noted
so the surviving selector is not "simplified" back into a silent
regression.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017YmbBhr8s7GusjDE12hrZM
2026-09-02 16:48:59 +01:00
TudorandClaude Opus 5 07d586d0ad docs(blog): how to publish, and why the app has two route groups
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m15s
PR Checks / Backend Smoke (pull_request) Successful in 10s
PR Checks / Build Backend (no push) (pull_request) Successful in 36s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m11s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 1m18s
PR Checks / AI Code Review (Claude) (pull_request) Failing after 2m43s
PUBLISHING.md carries the house style with the posts, so the standard
survives without the design doc to hand — including the rule that a post
states what a metric does not show, which is the strongest signal a
human wrote it.

CLAUDE.md gains the two constraints that are invisible from the code and
expensive to rediscover: metadata file conventions break if moved into a
route group, and the build must keep succeeding with DATABASE_URL unset.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017YmbBhr8s7GusjDE12hrZM
2026-09-02 16:30:09 +01:00
TudorandClaude Opus 5 b793640507 feat(blog): add the blog index, post pages, RSS and content sitemap
The rendering split is dictated by CI building with no database.

/blog, /blog/rss.xml and /content-sitemap.xml have no dynamic params, so
Next prerenders them at build time and the build fails on a missing
Payload secret — caught here, not on staging. They are force-dynamic
instead: one indexed query against Postgres on the same Docker network,
and a newly published post appears immediately rather than waiting on a
revalidation. /blog/[slug] keeps ISR, because with no
generateStaticParams there is nothing to prerender; it is generated on
first request and cached, which is exactly what the collection's
afterChange hook exists to invalidate.

RichText takes `converters`, not `blocks`, in Payload 3.88, and the
default converters must be spread or every paragraph and heading loses
its renderer and the body comes out empty.

BlogPosting references the Person and Organization by @id rather than
repeating them, so every post and the About page resolve to one author
entity instead of declaring several people with the same name.

/sitemap.xml is proxied from FastAPI, which knows nothing about Payload,
so the Next-owned URLs get their own sitemap and robots.txt lists both.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017YmbBhr8s7GusjDE12hrZM
2026-09-02 16:28:58 +01:00
TudorandClaude Opus 5 21a5d18f59 feat(blog): add the posts and media collections
Drafts are on so a post can be written across sittings without saving
being publishing.

afterChange and afterDelete revalidate every path a post appears on.
Blog pages are ISR because CI builds with no database, so without these
a published post would not appear until the revalidate window expired —
up to an hour of a writer concluding that publishing is broken. Payload
runs in the same process as Next, so these are direct revalidatePath
calls with no webhook and no shared secret.

Media writes to an absolute /app/media matching the compose mount; a
mismatch would write into the container filesystem, where the next
redeploy silently discards it. Alt text is required rather than
optional.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017YmbBhr8s7GusjDE12hrZM
2026-09-02 16:24:15 +01:00
TudorandClaude Opus 5 f614414070 feat(about): give the site a named author
The site had no author, no statement of why it exists and nobody
accountable for its numbers, which is most of why it reads as machine
generated.

The page states plainly that its author is not an education expert. The
credibility claim is lived experience — a parent going through primary
admissions — plus stated provenance for every figure, which is true and
cannot be undermined by someone noticing there is no teaching
qualification behind it. First name only: the Person JSON-LD carries no
familyName, worksFor or affiliation, and a test asserts it stays that
way.

The footer gains a fourth column, with a tablet breakpoint so four
columns pair up rather than crushing before the 768px collapse. The nav
is deliberately untouched — its mobile tab bar already carries four
items.

public/brand/tudor.jpg is NOT in this commit. The page references it and
will show a broken image until the photograph is supplied; a stock
portrait would defeat the entire point of the work.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017YmbBhr8s7GusjDE12hrZM
2026-09-02 16:22:52 +01:00
TudorandClaude Opus 5 310b63b0cb build(cms): wire Payload into the Docker image and both stacks
Uploads go to a named volume at /app/media. The directory is created in
the image before the mount and covered by the existing chown, because
Docker seeds a fresh named volume from the image path — a missing or
root-owned directory there fails every upload with EACCES at runtime,
long after the build passed.

PAYLOAD_SECRET uses the same :? form as AIRFLOW_ADMIN_PASSWORD: refuse
to start rather than boot with an empty secret and accept forged
sessions. Staging's must differ from production's, which the header
comment now says explicitly. Portainer prefixes volume names per stack,
so payload_media isolates itself.

prodMigrations is not wired yet — generating the initial migration needs
a reachable Postgres. Follows in its own commit.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017YmbBhr8s7GusjDE12hrZM
2026-09-02 16:19:27 +01:00
TudorandClaude Opus 5 c2c76c5817 feat(cms): keep the admin panel out of the index
X-Robots-Tag rather than the robots.txt Disallow alone, for the same
reason the staging rule uses one: a Disallow blocks crawling, not
indexing, so a URL found from an external link can be indexed without
ever being fetched — and blocking the crawl means the noindex is never
seen. Both mechanisms are applied to /admin and /cms-api.

The existing CSP is frame-ancestors only, which restricts who may embed
the site rather than what a page may load, so it cannot break the panel.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017YmbBhr8s7GusjDE12hrZM
2026-09-02 16:17:55 +01:00
TudorandClaude Opus 5 c5a4d106da feat(cms): install Payload and serve the admin panel
Payload 3.88 runs inside the Next app against the existing Postgres, in
its own 'payload' schema so no pipeline operation on public — the app
tables, Airflow's metadata, migrate_csv_to_db.py --drop — can reach blog
content.

Its REST API is mounted at /cms-api. /api is the FastAPI proxy's
catch-all, which would swallow every admin call and forward it to the
backend with no error. The mount points live in lib/payloadRoutes.ts so
there is one definition and a test can assert it without importing
Payload: it is ESM-only, next/jest will not transform it, and appending
transformIgnorePatterns cannot un-ignore a package. Forcing it through
transpilePackages would change how the production build bundles Payload
to serve a test, so the live proof that /api still reaches FastAPI stays
where it belongs — the e2e journeys, which call /api/schools.

The package becomes ESM ("type": "module"), which Payload's CLI requires:
richtext-lexical has top-level await and the config cannot be require()d.
Only two files needed renaming, jest.config.cjs and a build script.

The build is verified to succeed with DATABASE_URL and PAYLOAD_SECRET
both unset, which is how CI builds it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017YmbBhr8s7GusjDE12hrZM
2026-09-02 16:17:01 +01:00
TudorandClaude Opus 5 2437ffce42 refactor(app): move site routes into a (frontend) route group
Payload's admin panel ships its own root layout rendering html/body.
Next allows multiple root layouts only when no app/layout.tsx exists, so
the site's routes move into their own group. Route groups are invisible
to routing: every public URL is unchanged, verified against the build's
route table.

The metadata file conventions deliberately stay at the app/ root. Moving
them into the group renamed /icon.png to /icon-4usi79.png (likewise
apple-icon and opengraph-image) and dropped /robots.txt altogether,
which would have broken the /icon.png cache-control rule, the
outputFileTracingIncludes entry for the share card, and robots.txt.

darkThemeSafety reads app/globals.css off disk rather than importing it,
so it needed its own path fix — a grep for import specifiers misses it,
and it fails as an unrunnable suite rather than a failed assertion.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017YmbBhr8s7GusjDE12hrZM
2026-09-02 16:11:48 +01:00
TudorandClaude Opus 5 eb648f3f76 build(next): convert the config to ESM so Payload can wrap it
withPayload() is ESM-only, so next.config.js has to become .mjs. That
file also carries the rule that keeps staging out of Google's index, so
the conversion goes in on its own, behind a test that asserts the rule
survived — along with the standalone output, the opengraph-image font
tracing and the analytics frame-ancestors CSP.

Jest resolves the .mjs config without extra configuration, so
jest.config.js is untouched.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017YmbBhr8s7GusjDE12hrZM
2026-09-02 16:08:58 +01:00
TudorandClaude Opus 5 74e5fffc10 docs(about-blog): implementation plan for the About page and Payload blog
Nine tasks, each ending in an independently testable deliverable.

Two structural findings that the spec did not anticipate, both recorded
in the plan. Payload's admin panel ships its own root layout rendering
html/body, and Next allows multiple root layouts only when no
app/layout.tsx exists — so every existing route moves into an
app/(frontend) route group first, on its own, with the full suite as the
gate. Route groups are invisible to routing, so no public URL changes.

The second finding corrects the spec: adding /cms-api to the FastAPI
proxy's exclusion list would be dead code, because that catch-all only
ever matches /api/*. The route remap alone is sufficient.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017YmbBhr8s7GusjDE12hrZM
2026-09-02 15:59:44 +01:00
TudorandClaude Opus 5 748ef32180 docs(about-blog): design for a named author, an About page and a Payload blog
The site reads as synthetic because nobody is accountable for the
numbers, no editorial judgement is visible, and the voice is
institutional third person. This designs the fix: a named author
(first name, photo, explicitly not an education expert), a coded
/about page, and a blog backed by Payload CMS running inside the
existing Next app.

Also fills a hole in the SEO programme, which has eight workstreams
and no E-E-A-T or authorship signal on a YMYL corpus.

Records two collisions found while designing, both of which fail
badly if missed: Payload's default /api route fights the existing
FastAPI catch-all proxy, and withPayload() is ESM-only so
next.config.js — which carries staging's noindex header — has to
become next.config.mjs.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FT1Ls4GbgLDXoQX7NAuHGT
2026-09-02 14:43:20 +01:00
tudor b0c4ea8282 Merge pull request 'fix(destinations): the DAG died on a null in a primary key' (#139) from fix/destinations-national-grain into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 14s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 51s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 1m36s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Successful in 2m6s
Reviewed-on: #139
2026-08-31 20:43:26 +00:00
106 changed files with 13666 additions and 2028 deletions

No files matched your search

+13 -5
View File
@@ -20,7 +20,7 @@ PORT=80
# =============================================================================
# CORS
# =============================================================================
# Comma-separated list of allowed origins
# JSON array of allowed origins (pydantic-settings format)
# In production, only include your actual domain
ALLOWED_ORIGINS=["https://schoolcompare.co.uk"]
@@ -33,13 +33,21 @@ ADMIN_API_KEY=CHANGE_THIS_TO_A_SECURE_RANDOM_KEY
# Rate limiting (requests per minute per IP)
RATE_LIMIT_PER_MINUTE=60
RATE_LIMIT_BURST=10
GLOBAL_RATE_LIMIT_PER_MINUTE=3000
# Maximum request body size in bytes (default 1MB)
MAX_REQUEST_SIZE=1048576
# =============================================================================
# API
# SEARCH AND OPTIONAL FEATURE FLAGS
# =============================================================================
DEFAULT_PAGE_SIZE=50
MAX_PAGE_SIZE=100
TYPESENSE_URL=http://localhost:8108
TYPESENSE_API_KEY=CHANGE_THIS_TO_YOUR_TYPESENSE_KEY
# Empty URL disables Unleash-backed flags. Match the managed environment when used.
UNLEASH_URL=
UNLEASH_API_TOKEN=
# Page-size limits are currently declared by route Query parameters.
# DEFAULT_PAGE_SIZE, MAX_PAGE_SIZE and RATE_LIMIT_BURST are not reliable tuning
# controls in the current routes; see docs/LEGACY_CODE.md.
+13 -187
View File
@@ -1,191 +1,17 @@
# Docker Deployment Guide
# Docker deployment
## Quick Start
The maintained deployment runbook is [docs/DEPLOY.md](docs/DEPLOY.md).
Deploy the complete SchoolCompare stack (PostgreSQL + FastAPI + Next.js) with one command:
- Production: `docker-compose.portainer.yml`, using `:prod` images.
- Staging: `docker-compose.portainer.staging.yml`, using `:staging` images.
- Builds and deployment: `.gitea/workflows/deploy.yml`.
- Human-approved production promotion: `.gitea/workflows/promote.yml`.
```bash
docker-compose up -d
```
The generic `docker-compose.yml` is not a supported one-command onboarding path:
it still uses `:latest` tags that the release workflow no longer publishes and
lacks the full current CMS setup. Review the [legacy inventory](docs/LEGACY_CODE.md)
before using old compose examples. Starting an empty database does not populate
school marts.
This will start:
- **PostgreSQL** on port 5432 (database)
- **FastAPI** on port 8000 (backend API)
- **Next.js** on port 3000 (frontend)
## Service Details
### PostgreSQL Database
- **Port**: 5432
- **Container**: `schoolcompare_db`
- **Credentials**:
- User: `schoolcompare`
- Password: `schoolcompare`
- Database: `schoolcompare`
- **Volume**: `postgres_data` (persistent storage)
### FastAPI Backend
- **Port**: 8000 → 80 (container)
- **Container**: `schoolcompare_backend`
- **Built from**: Root `Dockerfile`
- **API Endpoint**: http://localhost:8000/api
- **Health Check**: http://localhost:8000/api/data-info
### Next.js Frontend
- **Port**: 3000
- **Container**: `schoolcompare_nextjs`
- **Built from**: `nextjs-app/Dockerfile`
- **URL**: http://localhost:3000
- **Connects to**: Backend via internal network
## Commands
### Start all services
```bash
docker-compose up -d
```
### View logs
```bash
# All services
docker-compose logs -f
# Specific service
docker-compose logs -f nextjs
docker-compose logs -f backend
docker-compose logs -f db
```
### Check status
```bash
docker-compose ps
```
### Stop all services
```bash
docker-compose down
```
### Rebuild after code changes
```bash
# Rebuild and restart specific service
docker-compose up -d --build nextjs
# Rebuild all services
docker-compose up -d --build
```
### Clean restart (remove volumes)
```bash
docker-compose down -v
docker-compose up -d
```
## Initial Database Setup
After first start, you may need to initialize the database:
```bash
# Enter the backend container
docker exec -it schoolcompare_backend bash
# Run migrations or data loading
python -m backend.data_loader
```
## Accessing Services
Once running:
- **Frontend**: http://localhost:3000
- **Backend API**: http://localhost:8000/api
- **API Docs**: http://localhost:8000/docs (Swagger UI)
- **Database**: localhost:5432 (use any PostgreSQL client)
## Environment Variables
Create a `.env` file in the root directory to customize:
```env
# Database
POSTGRES_USER=schoolcompare
POSTGRES_PASSWORD=your_secure_password
POSTGRES_DB=schoolcompare
# Backend
DATABASE_URL=postgresql://schoolcompare:your_secure_password@db:5432/schoolcompare
# Frontend (for client-side access)
NEXT_PUBLIC_API_URL=http://localhost:8000/api
```
Then run:
```bash
docker-compose up -d
```
## Troubleshooting
### Backend not connecting to database
```bash
# Check database health
docker-compose ps
# View backend logs
docker-compose logs backend
# Restart backend
docker-compose restart backend
```
### Frontend not connecting to backend
```bash
# Check backend health
curl http://localhost:8000/api/data-info
# Check Next.js environment variables
docker exec schoolcompare_nextjs env | grep API
```
### Port already in use
```bash
# Change ports in docker-compose.yml
# For example, change "3000:3000" to "3001:3000"
```
### Rebuild from scratch
```bash
docker-compose down -v
docker system prune -a
docker-compose up -d --build
```
## Production Deployment
For production, update the following:
1. **Use secure passwords** in `.env` file
2. **Configure reverse proxy** (Nginx) in front of Next.js
3. **Enable HTTPS** with SSL certificates
4. **Set production environment variables**:
```env
NODE_ENV=production
POSTGRES_PASSWORD=<strong-password>
```
5. **Backup database** regularly:
```bash
docker exec schoolcompare_db pg_dump -U schoolcompare schoolcompare > backup.sql
```
## Network Architecture
```
Internet
↓
Next.js (port 3000) ← User browsers
↓ (internal network)
FastAPI (port 8000) ← API calls
↓ (internal network)
PostgreSQL (port 5432) ← Data queries
```
All services communicate via the `schoolcompare-network` Docker network.
For architecture, configuration and test commands, see
[ARCHITECTURE.md](docs/ARCHITECTURE.md) and [DEVELOPMENT.md](docs/DEVELOPMENT.md).
+2
View File
@@ -1,3 +1,5 @@
> Historical migration record, retained for context. Setup and architecture claims below may be obsolete. Use [README.md](README.md), [architecture](docs/ARCHITECTURE.md) and [deployment](docs/DEPLOY.md) for current guidance.
# SchoolCompare: Vanilla JS → Next.js Migration Summary
## Overview
+52 -199
View File
@@ -1,214 +1,67 @@
# Primary School Compass 🧒📚
# SchoolCompare
A modern web application for comparing **primary school (KS2)** performance data in **Wandsworth and Merton** over the last 5 years. Built with FastAPI and vanilla JavaScript with Chart.js visualizations.
SchoolCompare compares schools across England: primary (KS2), secondary (KS4),
all-through and post-16 provision, with coverage depending on the source dataset.
It provides school search, postcode maps, comparisons, rankings, place pages,
Ofsted information, admissions and destination measures. Editorial content lives
in a Payload CMS blog.
![Python](https://img.shields.io/badge/Python-3.9+-blue)
![FastAPI](https://img.shields.io/badge/FastAPI-0.109-green)
![License](https://img.shields.io/badge/License-MIT-yellow)
## Start here
## Features
- [Architecture and data flow](docs/ARCHITECTURE.md)
- [Development and validation](docs/DEVELOPMENT.md)
- [Deployment and promotion](docs/DEPLOY.md)
- [Legacy and unused-code inventory](docs/LEGACY_CODE.md)
- [Frontend conventions](nextjs-app/README.md)
- [CMS publishing](nextjs-app/docs/PUBLISHING.md)
- 📊 **Interactive Charts** - Visualize KS2 performance trends over time
- 🔍 **Smart Search** - Find primary schools by name in Wandsworth & Merton
- ⚖️ **Side-by-Side Comparison** - Compare up to 5 schools simultaneously
- 🏆 **Rankings** - View top-performing primary schools by various KS2 metrics
- 📱 **Responsive Design** - Works beautifully on desktop and mobile
## Repository map
## Key Metrics (KS2)
| Path | Responsibility |
|---|---|
| `backend/` | FastAPI routes, cached school data, read-only SQLAlchemy mappings, feature flags |
| `nextjs-app/` | Next.js App Router, React UI, Payload CMS, frontend tests |
| `pipeline/plugins/extractors/` | Custom Singer taps for GIAS, EES, Ofsted and other datasets |
| `pipeline/transform/` | dbt staging/intermediate models, marts, seeds and data tests |
| `pipeline/dags/` | Airflow extraction, transformation and publication workflows |
| `pipeline/scripts/` | Search indexing, code generation and operational diagnostics |
| `e2e/` | Playwright journeys against a running environment |
| `.gitea/workflows/` | PR checks, staging deployment and manual production promotion |
| `scripts/` | CI review tooling and historical data utilities; see the legacy inventory |
| `docs/superpowers/`, `mockups/` | Design history and prototypes, not application entry points |
The application tracks these Key Stage 2 performance indicators:
## Runtime
| Metric | Description |
|--------|-------------|
| **Reading Progress** | Progress in reading from KS1 to KS2 |
| **Writing Progress** | Progress in writing from KS1 to KS2 |
| **Maths Progress** | Progress in maths from KS1 to KS2 |
| **Reading Expected %** | Percentage meeting expected standard in reading |
| **Writing Expected %** | Percentage meeting expected standard in writing |
| **Maths Expected %** | Percentage meeting expected standard in maths |
| **Reading, Writing & Maths Combined %** | Percentage meeting expected standard in all three subjects |
The public site is **Next.js**, not the FastAPI root page. Browser `/api/*`
requests pass through a Next.js route handler to FastAPI. Server-rendered pages
call FastAPI directly using `FASTAPI_URL`, including its `/api` suffix.
## Quick Start
PostgreSQL/PostGIS stores school data. Meltano/Singer extracts source data;
dbt builds `marts.*`; FastAPI reads those tables. Typesense serves text search
and autocomplete. Payload runs inside Next.js and owns a separate `payload`
database schema and uploaded media.
### 1. Clone and Setup
There is **no automatic CSV import or sample dataset on startup**. A working
school-data environment needs populated marts from the pipeline or an approved
database snapshot. See [development](docs/DEVELOPMENT.md) before choosing a setup.
```bash
cd school_results
## Validation
# Create virtual environment
python -m venv venv
source venv/bin/activate # On Windows: venv\Scripts\activate
# Install dependencies
pip install -r requirements.txt
```sh
cd nextjs-app
npm ci
npm run typecheck
npm test -- --runInBand
```
### 2. Run the Application
Backend checks, pipeline validation, runtime versions and E2E requirements are
listed in [DEVELOPMENT.md](docs/DEVELOPMENT.md). No `npm run lint` script is
currently defined.
```bash
# Start the server
python -m uvicorn backend.app:app --reload --port 8000
```
Then open http://localhost:8000 in your browser.
The app will run with **sample data** by default, showing **110 primary schools** (66 in Wandsworth, 44 in Merton) with 5 years of KS2 performance data.
### 3. (Optional) Use Real Data
To use real UK school performance data:
1. Visit [Compare School Performance - Download Data](https://www.compare-school-performance.service.gov.uk/download-data)
2. Download **Key Stage 2** data for the years you want (2019-2024)
- Select "Key Stage 2" as the data type
3. Place the CSV files in the `data/` folder
4. Restart the server - it will automatically load and filter to Wandsworth & Merton schools
**Note:** The app only displays schools in Wandsworth and Merton. Data from other areas will be filtered out.
See the helper script for more details:
```bash
python scripts/download_data.py
```
## Project Structure
```
school_results/
├── backend/
│ └── app.py # FastAPI application with all API endpoints
├── frontend/
│ ├── index.html # Main HTML page
│ ├── styles.css # Styling (warm, editorial design)
│ └── app.js # Frontend JavaScript
├── data/
│ └── .gitkeep # Place CSV data files here
├── scripts/
│ └── download_data.py # Helper for downloading/processing data
├── requirements.txt # Python dependencies
└── README.md
```
## API Endpoints
| Endpoint | Description |
|----------|-------------|
| `GET /api/schools` | List schools with optional search/filter |
| `GET /api/schools/{urn}` | Get detailed data for a specific school |
| `GET /api/compare?urns=...` | Compare multiple schools |
| `GET /api/rankings` | Get school rankings by metric |
| `GET /api/filters` | Get available filter options |
| `GET /api/metrics` | Get available performance metrics |
### Example API Usage
```bash
# Search for schools
curl "http://localhost:8000/api/schools?search=academy"
# Get school details
curl "http://localhost:8000/api/schools/100001"
# Compare schools
curl "http://localhost:8000/api/compare?urns=100001,100002,100003"
# Get rankings
curl "http://localhost:8000/api/rankings?metric=rwm_expected_pct&year=2024"
```
## Data Format
If using your own CSV data, ensure it includes these columns (or similar):
| Column | Type | Description |
|--------|------|-------------|
| URN | Integer | Unique Reference Number |
| SCHNAME | String | School name |
| LA | String | Local Authority (must be Wandsworth or Merton) |
| READPROG | Float | Reading progress score |
| WRITPROG | Float | Writing progress score |
| MATPROG | Float | Maths progress score |
| PTRWM_EXP | Float | % meeting expected standard in reading, writing & maths |
| PTREAD_EXP | Float | % meeting expected standard in reading |
| PTWRIT_EXP | Float | % meeting expected standard in writing |
| PTMAT_EXP | Float | % meeting expected standard in maths |
The application normalizes column names automatically and filters to only show Wandsworth and Merton schools.
## Technology Stack
- **Backend**: FastAPI (Python) - High-performance async API framework
- **Frontend**: Vanilla JavaScript with Chart.js
- **Styling**: Custom CSS with CSS variables for theming
- **Data**: Pandas for CSV processing
## Design Philosophy
The UI features a warm, editorial design inspired by quality publications:
- **Typography**: DM Sans for body text, Playfair Display for headings
- **Color Palette**: Warm cream background with coral and teal accents
- **Interactions**: Smooth animations and hover effects
- **Charts**: Clean, readable data visualizations
## Development
```bash
# Run with auto-reload
python -m uvicorn backend.app:app --reload --port 8000
# Or run directly
python backend/app.py
```
## Coverage
This application is specifically designed for:
- **School Phase**: Primary schools only (Key Stage 2)
- **Geographic Area**: Wandsworth and Merton (London boroughs)
- **Time Period**: Last 5 years of data (2020-2024)
Note: 2021 data shows as unavailable because SATs were cancelled due to COVID-19.
## Data Source
Data is sourced from the UK Government's [Compare School Performance](https://www.compare-school-performance.service.gov.uk/) service, which provides official school performance data for England.
**Important**: When using real data, please comply with the [terms of use](https://www.compare-school-performance.service.gov.uk/download-data) and data protection regulations.
## Scheduled Jobs
### Geocoding Schools (Cron Job)
School postcodes are geocoded by a scheduled job, not on-demand. This improves performance and reduces API calls.
**Setup the cron job** (runs weekly on Sunday at 2am):
```bash
# Edit crontab
crontab -e
# Add this line (adjust paths as needed):
0 2 * * 0 cd /path/to/school_compare && /path/to/venv/bin/python scripts/geocode_schools.py >> /var/log/geocode_schools.log 2>&1
```
**Manual run:**
```bash
# Geocode only schools missing coordinates
python scripts/geocode_schools.py
# Force re-geocode all schools
python scripts/geocode_schools.py --force
```
## License
MIT License - feel free to use this project for educational purposes.
---
Built with ❤️ for Wandsworth & Merton families
## Deployment
Work on a feature branch and open a PR. Merging to `main` builds images and
deploys staging. Production promotion is a separate, human-triggered Gitea
workflow. Use [DEPLOY.md](docs/DEPLOY.md) and the Portainer compose files as the
operational references. The generic compose examples still reference `:latest`,
which the current release workflow does not publish.
+69 -1
View File
@@ -38,7 +38,7 @@ from .data_loader import (
)
from .data_loader import get_data_info as get_db_info
from . import flags
from .places import build_place_registry
from .places import build_place_index, build_place_registry, places_for_urn
from .schemas import METRIC_DEFINITIONS, RANKING_COLUMNS, SCHOOL_COLUMNS
from .utils import clean_for_json, convert_to_native
@@ -65,6 +65,10 @@ _sitemaps: dict[str, str] | None = None
# Built from the same DataFrame the sitemap uses, so places and sitemap can
# never describe different corpora. Reset by the same admin endpoint.
_place_registry: dict | None = None
# Cached beside the registry, and invalidated by identity against it — see
# get_place_index. Never cleared independently.
_place_index: dict | None = None
_place_index_source: dict | None = None
VALID_PLACE_KINDS = ("town", "locality", "authority", "outcode")
@@ -188,6 +192,24 @@ def get_place_registry() -> dict:
return _place_registry
def get_place_index() -> dict:
"""URN → its published places, cached against the registry it came from.
Invalidation is an identity check rather than a second flag to remember to
clear. Anything that drops `_place_registry` — the tests all do — gets a
fresh registry object here, which no longer matches the one the index was
built from, so the index rebuilds with it. A separate `_place_index = None`
would be one more thing to forget, and a stale reverse index is exactly the
bug that would put links to another dataset's places on a school page.
"""
global _place_index, _place_index_source
registry = get_place_registry()
if _place_index is None or _place_index_source is not registry:
_place_index = build_place_index(registry)
_place_index_source = registry
return _place_index
def _urlset(rows: list[str]) -> str:
return "\n".join([
'<?xml version="1.0" encoding="UTF-8"?>',
@@ -211,6 +233,45 @@ def _place_url(place) -> str:
return f"/schools/{place.slug}"
def _places_payload(urn: int) -> list[dict]:
"""The published places containing this school, as the school page needs
them: a name to write in the link, a count so the anchor can say what it
leads to, and the canonical path.
`phases` carries the phase variants this school actually appears on, which
is usually one and is two for an all-through school — it is listed on both
pages, so there is no tie to break.
Membership is read straight from the registry's own `phase_urns` rather
than re-derived from the school's phase string. The registry is the one
place that decides which phases a place publishes and who is on them;
computing it a second time here is how a page comes to link a school to a
phase page that does not list it, or to a route that does not exist. That
is also why outcodes need no special case: they carry empty `phase_urns`,
so they report no phase links on their own.
"""
payload = []
for place in places_for_urn(get_place_index(), int(urn)):
phases = [
{
"phase": phase,
"count": len(phase_urns),
"url": f"{_place_url(place)}/{phase}",
}
for phase, phase_urns in sorted(place.phase_urns.items())
if int(urn) in phase_urns
]
payload.append({
"kind": place.kind,
"slug": place.slug,
"name": place.name,
"count": len(place.urns),
"url": _place_url(place),
"phases": phases,
})
return payload
def _place_sitemap_rows(kinds: tuple[str, ...]) -> list[str]:
"""A <url> per place, plus a phase variant wherever that phase clears the
threshold on its own.
@@ -902,6 +963,13 @@ async def get_school_details(request: Request, urn: int):
return {
"school_info": school_info,
# Where this school sits in the location layer, for the page's link
# module and breadcrumb. Derived from the same registry the place
# pages and the sitemap use, so a link is never offered for a page
# that does not exist. Empty is a valid answer: a school whose town
# and authority both fall below the publish threshold has nowhere to
# point, and the page renders without the module.
"places": _places_payload(urn),
"yearly_data": clean_for_json(school_data),
# Supplementary data (null if not yet populated by Kestra)
"ofsted": supplementary.get("ofsted"),
-10
View File
@@ -188,16 +188,6 @@ def geocode_single_postcode(postcode: str) -> Optional[Tuple[float, float]]:
return None
def haversine_distance(lat1: float, lon1: float, lat2: float, lon2: float) -> float:
"""Calculate great-circle distance between two points (miles)."""
from math import radians, cos, sin, asin, sqrt
lat1, lon1, lat2, lon2 = map(radians, [lat1, lon1, lat2, lon2])
dlat = lat2 - lat1
dlon = lon2 - lon1
a = sin(dlat / 2) ** 2 + cos(lat1) * cos(lat2) * sin(dlon / 2) ** 2
return 2 * asin(sqrt(a)) * 3956
# =============================================================================
# MAIN DATA LOAD — joins dim_school + dim_location + fact_performance
# fact_performance is a merged KS2+KS4 table (one row per URN per year).
+17
View File
@@ -57,6 +57,23 @@ REGISTRY: dict[str, Flag] = {
),
added=date(2026, 8, 26),
),
Flag(
name="about_page",
description=(
"The /about page, its footer link, its sitemap entry, and the "
"named-author byline on every blog post."
),
added=date(2026, 9, 8),
),
Flag(
name="blog",
description=(
"The /blog index, post pages, the RSS feed, their footer link "
"and their sitemap entries. Not /admin: posts must be "
"writable before the blog is readable."
),
added=date(2026, 9, 8),
),
)
}
-512
View File
@@ -1,512 +0,0 @@
"""
Database migration logic for importing CSV data.
Used by both CLI script and automatic startup migration.
"""
import re
from pathlib import Path
from typing import Dict, Optional
import numpy as np
import pandas as pd
import requests
from .config import settings
from .database import Base, engine, get_db_session
from .models import School, SchoolResult
from .schemas import (
COLUMN_MAPPINGS,
LA_CODE_TO_NAME,
NULL_VALUES,
SCHOOL_TYPE_MAP,
)
def parse_numeric(value) -> Optional[float]:
"""Parse a numeric value, handling special cases."""
if pd.isna(value):
return None
if isinstance(value, (int, float)):
return float(value) if not np.isnan(value) else None
str_val = str(value).strip().upper()
if str_val in NULL_VALUES or str_val == "":
return None
# Remove percentage signs if present
str_val = str_val.replace("%", "")
try:
return float(str_val)
except ValueError:
return None
def extract_year_from_folder(folder_name: str) -> Optional[int]:
"""Extract year from folder name like '2023-2024'."""
match = re.search(r"(\d{4})-(\d{4})", folder_name)
if match:
return int(match.group(2))
match = re.search(r"(\d{4})", folder_name)
if match:
return int(match.group(1))
return None
def geocode_postcodes_bulk(postcodes: list) -> Dict[str, tuple]:
"""
Geocode postcodes in bulk using postcodes.io API.
Returns dict of postcode -> (latitude, longitude).
"""
results = {}
valid_postcodes = [
p.strip().upper()
for p in postcodes
if p and isinstance(p, str) and len(p.strip()) >= 5
]
valid_postcodes = list(set(valid_postcodes))
if not valid_postcodes:
return results
batch_size = 100
total_batches = (len(valid_postcodes) + batch_size - 1) // batch_size
for i, batch_start in enumerate(range(0, len(valid_postcodes), batch_size)):
batch = valid_postcodes[batch_start : batch_start + batch_size]
print(
f" Geocoding batch {i + 1}/{total_batches} ({len(batch)} postcodes)..."
)
try:
response = requests.post(
"https://api.postcodes.io/postcodes",
json={"postcodes": batch},
timeout=30,
)
if response.status_code == 200:
data = response.json()
for item in data.get("result", []):
if item and item.get("result"):
pc = item["query"].upper()
lat = item["result"].get("latitude")
lon = item["result"].get("longitude")
if lat and lon:
results[pc] = (lat, lon)
except Exception as e:
print(f" Warning: Geocoding batch failed: {e}")
return results
def load_csv_data(data_dir: Path) -> pd.DataFrame:
"""Load all CSV data from data directory."""
all_data = []
for folder in sorted(data_dir.iterdir()):
if not folder.is_dir():
continue
year = extract_year_from_folder(folder.name)
if not year:
continue
# Specifically look for the KS2 results file
ks2_file = folder / "england_ks2final.csv"
if not ks2_file.exists():
continue
csv_file = ks2_file
print(f" Loading {csv_file.name} (year {year})...")
try:
df = pd.read_csv(csv_file, encoding="latin-1", low_memory=False)
except Exception as e:
print(f" Error loading {csv_file}: {e}")
continue
# Rename columns
df.rename(columns=COLUMN_MAPPINGS, inplace=True)
df["year"] = year
# Handle local authority name
la_name_cols = ["LANAME", "LA (name)", "LA_NAME", "LA NAME"]
la_name_col = next((c for c in la_name_cols if c in df.columns), None)
if la_name_col and la_name_col != "local_authority":
df["local_authority"] = df[la_name_col]
elif "LEA" in df.columns:
df["local_authority_code"] = pd.to_numeric(df["LEA"], errors="coerce")
df["local_authority"] = (
df["local_authority_code"]
.map(LA_CODE_TO_NAME)
.fillna(df["LEA"].astype(str))
)
# Store LEA code
if "LEA" in df.columns:
df["local_authority_code"] = pd.to_numeric(df["LEA"], errors="coerce")
# Map school type
if "school_type_code" in df.columns:
df["school_type"] = (
df["school_type_code"]
.map(SCHOOL_TYPE_MAP)
.fillna(df["school_type_code"])
)
# Create combined address
addr_parts = ["address1", "address2", "town", "postcode"]
for col in addr_parts:
if col not in df.columns:
df[col] = None
df["address"] = df.apply(
lambda r: ", ".join(
str(v)
for v in [
r.get("address1"),
r.get("address2"),
r.get("town"),
r.get("postcode"),
]
if pd.notna(v) and str(v).strip()
),
axis=1,
)
all_data.append(df)
print(f" Loaded {len(df)} records")
if all_data:
result = pd.concat(all_data, ignore_index=True)
print(f"\nTotal records loaded: {len(result)}")
print(f"Unique schools: {result['urn'].nunique()}")
print(f"Years: {sorted(result['year'].unique())}")
return result
return pd.DataFrame()
def migrate_data(df: pd.DataFrame, geocode: bool = False, geocode_cache: dict = None):
"""Migrate DataFrame data to database."""
if geocode_cache is None:
geocode_cache = {}
# Clean URN column - convert to integer, drop invalid values
df = df.copy()
df["urn"] = pd.to_numeric(df["urn"], errors="coerce")
df = df.dropna(subset=["urn"])
df["urn"] = df["urn"].astype(int)
# Group by URN to get unique schools (use latest year's data)
school_data = (
df.sort_values("year", ascending=False).groupby("urn").first().reset_index()
)
print(f"\nMigrating {len(school_data)} unique schools...")
# Geocode postcodes that aren't already in the cache
geocoded = dict(geocode_cache) # start with preserved coordinates
if geocode and "postcode" in df.columns:
cached_postcodes = {
str(row.get("postcode", "")).strip().upper()
for _, row in school_data.iterrows()
if int(float(str(row.get("urn", 0) or 0))) in geocode_cache
}
postcodes_needed = [
p for p in df["postcode"].dropna().unique()
if str(p).strip().upper() not in cached_postcodes
]
if postcodes_needed:
print(f"\nGeocoding {len(postcodes_needed)} postcodes ({len(geocode_cache)} restored from cache)...")
fresh = geocode_postcodes_bulk(postcodes_needed)
geocoded.update(fresh)
print(f" Successfully geocoded {len(fresh)} new postcodes")
else:
print(f"\nAll {len(geocode_cache)} postcodes restored from cache, skipping geocoding.")
with get_db_session() as db:
# Create schools
urn_to_school_id = {}
schools_created = 0
for _, row in school_data.iterrows():
# Safely parse URN - handle None, NaN, whitespace, and invalid values
urn_val = row.get("urn")
urn = None
if pd.notna(urn_val):
try:
urn_str = str(urn_val).strip()
if urn_str:
urn = int(float(urn_str)) # Handle "12345.0" format
except (ValueError, TypeError):
pass
if not urn:
continue
# Skip if we've already added this URN (handles duplicates in source data)
if urn in urn_to_school_id:
continue
# Get geocoding data
postcode = row.get("postcode")
lat, lon = None, None
if postcode and pd.notna(postcode):
coords = geocoded.get(str(postcode).strip().upper())
if coords:
lat, lon = coords
# Safely parse local_authority_code
la_code = None
la_code_val = row.get("local_authority_code")
if pd.notna(la_code_val):
try:
la_code_str = str(la_code_val).strip()
if la_code_str:
la_code = int(float(la_code_str))
except (ValueError, TypeError):
pass
school = School(
urn=urn,
school_name=row.get("school_name")
if pd.notna(row.get("school_name"))
else "Unknown",
local_authority=row.get("local_authority")
if pd.notna(row.get("local_authority"))
else None,
local_authority_code=la_code,
school_type=row.get("school_type")
if pd.notna(row.get("school_type"))
else None,
school_type_code=row.get("school_type_code")
if pd.notna(row.get("school_type_code"))
else None,
religious_denomination=row.get("religious_denomination")
if pd.notna(row.get("religious_denomination"))
else None,
age_range=row.get("age_range")
if pd.notna(row.get("age_range"))
else None,
address1=row.get("address1") if pd.notna(row.get("address1")) else None,
address2=row.get("address2") if pd.notna(row.get("address2")) else None,
town=row.get("town") if pd.notna(row.get("town")) else None,
postcode=row.get("postcode") if pd.notna(row.get("postcode")) else None,
latitude=lat,
longitude=lon,
)
db.add(school)
db.flush() # Get the ID
urn_to_school_id[urn] = school.id
schools_created += 1
if schools_created % 1000 == 0:
print(f" Created {schools_created} schools...")
print(f" Created {schools_created} schools")
# Create results
print(f"\nMigrating {len(df)} yearly results...")
results_created = 0
for _, row in df.iterrows():
# Safely parse URN
urn_val = row.get("urn")
urn = None
if pd.notna(urn_val):
try:
urn_str = str(urn_val).strip()
if urn_str:
urn = int(float(urn_str))
except (ValueError, TypeError):
pass
if not urn or urn not in urn_to_school_id:
continue
school_id = urn_to_school_id[urn]
# Safely parse year
year_val = row.get("year")
year = None
if pd.notna(year_val):
try:
year = int(float(str(year_val).strip()))
except (ValueError, TypeError):
pass
if not year:
continue
result = SchoolResult(
school_id=school_id,
year=year,
total_pupils=parse_numeric(row.get("total_pupils")),
eligible_pupils=parse_numeric(row.get("eligible_pupils")),
# Expected Standard
rwm_expected_pct=parse_numeric(row.get("rwm_expected_pct")),
reading_expected_pct=parse_numeric(row.get("reading_expected_pct")),
writing_expected_pct=parse_numeric(row.get("writing_expected_pct")),
maths_expected_pct=parse_numeric(row.get("maths_expected_pct")),
gps_expected_pct=parse_numeric(row.get("gps_expected_pct")),
science_expected_pct=parse_numeric(row.get("science_expected_pct")),
# Higher Standard
rwm_high_pct=parse_numeric(row.get("rwm_high_pct")),
reading_high_pct=parse_numeric(row.get("reading_high_pct")),
writing_high_pct=parse_numeric(row.get("writing_high_pct")),
maths_high_pct=parse_numeric(row.get("maths_high_pct")),
gps_high_pct=parse_numeric(row.get("gps_high_pct")),
# Progress
reading_progress=parse_numeric(row.get("reading_progress")),
writing_progress=parse_numeric(row.get("writing_progress")),
maths_progress=parse_numeric(row.get("maths_progress")),
# Averages
reading_avg_score=parse_numeric(row.get("reading_avg_score")),
maths_avg_score=parse_numeric(row.get("maths_avg_score")),
gps_avg_score=parse_numeric(row.get("gps_avg_score")),
# Context
disadvantaged_pct=parse_numeric(row.get("disadvantaged_pct")),
eal_pct=parse_numeric(row.get("eal_pct")),
sen_support_pct=parse_numeric(row.get("sen_support_pct")),
sen_ehcp_pct=parse_numeric(row.get("sen_ehcp_pct")),
stability_pct=parse_numeric(row.get("stability_pct")),
# Absence
reading_absence_pct=parse_numeric(row.get("reading_absence_pct")),
gps_absence_pct=parse_numeric(row.get("gps_absence_pct")),
maths_absence_pct=parse_numeric(row.get("maths_absence_pct")),
writing_absence_pct=parse_numeric(row.get("writing_absence_pct")),
science_absence_pct=parse_numeric(row.get("science_absence_pct")),
# Gender
rwm_expected_boys_pct=parse_numeric(row.get("rwm_expected_boys_pct")),
rwm_expected_girls_pct=parse_numeric(row.get("rwm_expected_girls_pct")),
rwm_high_boys_pct=parse_numeric(row.get("rwm_high_boys_pct")),
rwm_high_girls_pct=parse_numeric(row.get("rwm_high_girls_pct")),
# Disadvantaged
rwm_expected_disadvantaged_pct=parse_numeric(
row.get("rwm_expected_disadvantaged_pct")
),
rwm_expected_non_disadvantaged_pct=parse_numeric(
row.get("rwm_expected_non_disadvantaged_pct")
),
disadvantaged_gap=parse_numeric(row.get("disadvantaged_gap")),
# 3-Year
rwm_expected_3yr_pct=parse_numeric(row.get("rwm_expected_3yr_pct")),
reading_avg_3yr=parse_numeric(row.get("reading_avg_3yr")),
maths_avg_3yr=parse_numeric(row.get("maths_avg_3yr")),
)
db.add(result)
results_created += 1
if results_created % 10000 == 0:
print(f" Created {results_created} results...")
db.flush()
print(f" Created {results_created} results")
# Commit all changes
db.commit()
print("\nMigration complete!")
def _apply_schema_alterations():
"""
Add new columns to existing tables using ALTER TABLE … ADD COLUMN IF NOT EXISTS.
Safe to run on every migration — no-ops if the column already exists.
Add entries here whenever models.py gains new columns on an existing table.
"""
alterations = [
# v4: Ofsted Report Card columns
"ALTER TABLE ofsted_inspections ADD COLUMN IF NOT EXISTS framework VARCHAR(20)",
"ALTER TABLE ofsted_inspections ADD COLUMN IF NOT EXISTS rc_safeguarding_met BOOLEAN",
"ALTER TABLE ofsted_inspections ADD COLUMN IF NOT EXISTS rc_inclusion INTEGER",
"ALTER TABLE ofsted_inspections ADD COLUMN IF NOT EXISTS rc_curriculum_teaching INTEGER",
"ALTER TABLE ofsted_inspections ADD COLUMN IF NOT EXISTS rc_achievement INTEGER",
"ALTER TABLE ofsted_inspections ADD COLUMN IF NOT EXISTS rc_attendance_behaviour INTEGER",
"ALTER TABLE ofsted_inspections ADD COLUMN IF NOT EXISTS rc_personal_development INTEGER",
"ALTER TABLE ofsted_inspections ADD COLUMN IF NOT EXISTS rc_leadership_governance INTEGER",
"ALTER TABLE ofsted_inspections ADD COLUMN IF NOT EXISTS rc_early_years INTEGER",
"ALTER TABLE ofsted_inspections ADD COLUMN IF NOT EXISTS rc_sixth_form INTEGER",
]
from sqlalchemy import text as sa_text
with engine.connect() as conn:
for stmt in alterations:
try:
conn.execute(sa_text(stmt))
except Exception as e:
print(f" Warning: alteration skipped ({e})")
conn.commit()
def _apply_schema_drops():
"""
Drop tables retired from the schema. Idempotent (DROP … IF EXISTS), so it's
safe to run on every migration. Add entries here when a model is removed.
"""
drops = [
# v6: Ofsted Parent View feature removed
"DROP TABLE IF EXISTS marts.fact_parent_view CASCADE",
]
from sqlalchemy import text as sa_text
with engine.connect() as conn:
for stmt in drops:
try:
conn.execute(sa_text(stmt))
except Exception as e:
print(f" Warning: drop skipped ({e})")
conn.commit()
def run_full_migration(geocode: bool = False) -> bool:
"""
Run a complete migration: drop all tables and reimport from CSV.
Returns True if successful, False if no data found.
Raises exception on error.
"""
# Preserve existing geocoding so a reimport doesn't throw away coordinates
# that took a long time to compute.
geocode_cache: dict[int, tuple[float, float]] = {}
inspector = __import__("sqlalchemy").inspect(engine)
if "schools" in inspector.get_table_names():
try:
with get_db_session() as db:
rows = db.execute(
__import__("sqlalchemy").text(
"SELECT urn, latitude, longitude FROM schools "
"WHERE latitude IS NOT NULL AND longitude IS NOT NULL"
)
).fetchall()
geocode_cache = {r.urn: (r.latitude, r.longitude) for r in rows}
print(f" Saved {len(geocode_cache)} existing geocoded coordinates.")
except Exception as e:
print(f" Warning: could not save geocode cache: {e}")
# Only drop the core KS2 tables — leave supplementary tables (ofsted, census,
# finance, etc.) intact so a reimport doesn't wipe integrator-populated data.
# schema_version is NOT dropped: it persists so restarts don't re-trigger migration.
ks2_tables = ["school_results", "schools"]
print(f"Dropping core tables: {ks2_tables} ...")
inspector = __import__("sqlalchemy").inspect(engine)
existing = set(inspector.get_table_names())
for tname in ks2_tables:
if tname in existing:
Base.metadata.tables[tname].drop(bind=engine)
print("Creating all tables...")
Base.metadata.create_all(bind=engine)
# ALTER existing supplementary tables to add any new columns.
# create_all() only creates missing tables; it won't add columns to tables
# that already exist from an older schema version. These statements are
# idempotent (IF NOT EXISTS) so they're safe to run on every migration.
print("Applying column additions to supplementary tables...")
_apply_schema_alterations()
print("Dropping retired tables...")
_apply_schema_drops()
print("\nLoading CSV data...")
df = load_csv_data(settings.data_dir)
if df.empty:
print("Warning: No CSV data found to migrate!")
return False
migrate_data(df, geocode=geocode, geocode_cache=geocode_cache)
return True
+42
View File
@@ -296,6 +296,48 @@ def _locality_places(df, publishable: set[int],
return out
# Ordered authority → town/locality → outcode, widest first, because that is
# the order a breadcrumb reads. The link module re-sorts for its own purposes.
_PLACE_ORDER = {"authority": 0, "town": 1, "locality": 2, "outcode": 3}
def build_place_index(registry: dict[str, Place]) -> dict[int, tuple[Place, ...]]:
"""URN → the published places containing it, built once per registry.
The reverse of the registry, and the thing school pages link out through.
Derived from the registry rather than maintained beside it, so the two
cannot disagree about which places exist: a place below the publish
threshold is absent from the registry, so it is absent from here too, and
a link is never offered for a page that does not exist.
Built as an index rather than scanned per call because /api/schools/{urn}
is the site's highest-traffic endpoint. Scanning meant walking every place
and doing a tuple membership test against each — on the order of 10^5
comparisons per request, repeated for every school page view. One pass at
registry-build time replaces all of it with a dict lookup.
"""
grouped: dict[int, list[Place]] = {}
for place in registry.values():
for urn in place.urns:
grouped.setdefault(int(urn), []).append(place)
return {
urn: tuple(sorted(places,
key=lambda p: (_PLACE_ORDER.get(p.kind, 9), p.slug)))
for urn, places in grouped.items()
}
def places_for_urn(index: dict[int, tuple[Place, ...]], urn: int) -> tuple[Place, ...]:
"""The published places containing this school, widest first.
Empty is a real answer, not a failure: a school whose town and authority
both fall below the publish threshold has nowhere to link, and the page
renders without the module.
"""
return index.get(int(urn), ())
def build_place_registry(df) -> dict[str, Place]:
"""Every place the site publishes, keyed by "<kind>:<slug>"."""
if df.empty or "urn" not in df.columns:
+78 -1
View File
@@ -8,7 +8,8 @@ import numpy as np
import pandas as pd
import pytest
from backend.places import MIN_SCHOOLS, build_place_registry
from backend.places import (MIN_SCHOOLS, build_place_index,
build_place_registry, places_for_urn)
def _df(rows: list[dict]) -> pd.DataFrame:
@@ -418,3 +419,79 @@ def test_an_authority_still_publishes_phase_variants():
and /schools/authority/[la]/[phase] is the route that serves it."""
reg = build_place_registry(_df(_town(MIN_SCHOOLS, "Maidstone", "Kent")))
assert reg["authority:kent"].publishes_phase("primary")
# ── The reverse index: which published places contain a school ──────────────
#
# School pages link out to the location layer through this. It is the whole
# point of the index: before it, ~27k school pages linked to nothing on the
# site and stranded whatever authority they held.
def test_a_school_resolves_to_every_published_place_containing_it():
reg = build_place_registry(_df(_town(MIN_SCHOOLS, "Brentwood", "Essex")))
places = places_for_urn(build_place_index(reg), 100000)
kinds = {p.kind for p in places}
assert "town" in kinds
assert "authority" in kinds
def test_a_school_in_an_unpublished_town_still_resolves_to_its_authority():
# A town below the threshold has no page, so there is no link to offer —
# but the authority above it clears the threshold on the same schools and
# is where that reader should be sent.
reg = build_place_registry(_df(
_town(MIN_SCHOOLS - 1, "Tinytown", "Essex")
+ _town(MIN_SCHOOLS, "Brentwood", "Essex", start=200000)
))
places = places_for_urn(build_place_index(reg), 100000)
# The town is below the threshold, so it has no page and must not be
# offered as a link. The authority above it does, and is the right target.
assert all(p.slug != "tinytown" for p in places)
assert "authority" in {p.kind for p in places}
def test_an_unknown_urn_resolves_to_nothing_rather_than_raising():
# A school page renders for any URN the API knows; the link module is not
# entitled to take the page down when it has nothing to say.
reg = build_place_registry(_df(_town(MIN_SCHOOLS, "Brentwood", "Essex")))
assert places_for_urn(build_place_index(reg), 999999) == ()
def test_the_index_is_consistent_with_the_registry_it_was_built_from():
# The invariant that matters: a link module must never offer a place whose
# page does not exist, and never omit one that does.
reg = build_place_registry(_df(
_town(MIN_SCHOOLS, "Brentwood", "Essex")
+ _town(MIN_SCHOOLS, "Bedford", "Bedford", start=300000)
))
index = build_place_index(reg)
for key, place in reg.items():
for urn in place.urns:
assert place in places_for_urn(index, urn), (
f"{urn} is in {key} but the index does not say so")
def test_the_index_holds_no_school_the_registry_does_not():
# The reverse direction of the invariant above. An index entry for a URN
# no published place contains would put a link on a page for a place that
# does not list that school.
reg = build_place_registry(_df(
_town(MIN_SCHOOLS, "Brentwood", "Essex")
+ _town(MIN_SCHOOLS - 1, "Tinytown", "Essex", start=400000)
))
index = build_place_index(reg)
for urn, places in index.items():
for place in places:
assert urn in place.urns
assert place.key in reg
def test_the_index_preserves_the_widest_first_order():
# The breadcrumb reads authority then town, and takes this order as given.
reg = build_place_registry(_df(_town(MIN_SCHOOLS, "Brentwood", "Essex")))
kinds = [p.kind for p in places_for_urn(build_place_index(reg), 100000)]
assert kinds.index("authority") < kinds.index("town")
+176
View File
@@ -56,6 +56,11 @@ def client(monkeypatch):
monkeypatch.setattr(
app_module, "get_supplementary_data", lambda db, urn: {}
)
# The place registry is a module-level cache, so without this the endpoint
# answers from whatever registry an earlier test happened to leave behind
# — and a `places == []` assertion is satisfied by a stale registry just
# as well as by this fixture's own data, which makes it prove nothing.
monkeypatch.setattr(app_module, "_place_registry", None)
return TestClient(app_module.app, raise_server_exceptions=False)
@@ -69,3 +74,174 @@ def test_nan_gias_fields_serialize_as_null(client):
assert info["capacity"] is None
assert info["total_pupils"] is None
assert info["school_name"] == "West London Performing Arts Academy"
# ── Links out to the location layer ─────────────────────────────────────────
#
# School pages carried no link into the site at all: the only anchor on the
# template pointed at the school's own website, so ~27k pages received
# whatever authority the site had and sent it off-site. `places` is what the
# link module and the breadcrumb are built from.
def test_places_is_present_even_when_the_school_belongs_to_none(client):
# This fixture's single school cannot clear any publish threshold, so the
# honest answer is an empty list. The key must still be there: a missing
# key and "no places" are different things to the page rendering it.
body = client.get("/api/schools/150275").json()
assert body["places"] == []
def test_places_names_only_pages_that_exist(monkeypatch):
from backend import app as app_module
from backend.places import MIN_SCHOOLS
def _df():
return pd.DataFrame([
{
"urn": 100000 + i,
"school_name": f"Brentwood School {i}",
"town": "Brentwood",
"local_authority": "Essex",
"postcode": "CM15 8AA",
"phase": "Primary",
"year": 202425,
"rwm_expected_pct": 60.0,
"attainment_8_score": np.nan,
"ofsted_grade": 2.0,
"ofsted_date": None,
}
for i in range(MIN_SCHOOLS)
])
monkeypatch.setattr(app_module, "load_school_data", _df)
monkeypatch.setattr(app_module, "get_supplementary_data", lambda db, urn: {})
monkeypatch.setattr(app_module, "_place_registry", None)
client = TestClient(app_module.app, raise_server_exceptions=False)
places = client.get("/api/schools/100000").json()["places"]
assert places, "a school in a published town must offer links"
by_kind = {p["kind"]: p for p in places}
assert by_kind["town"]["url"] == "/schools/brentwood"
assert by_kind["authority"]["url"] == "/schools/authority/essex"
# Every entry carries what the link text needs, and a count, so the anchor
# can say what it leads to rather than "click here".
for place in places:
assert place["name"]
assert place["count"] >= 1
assert place["url"].startswith("/schools/")
def _brentwood_df(phase: str = "Primary", n: int = None):
from backend.places import MIN_SCHOOLS
n = n if n is not None else MIN_SCHOOLS
return lambda: pd.DataFrame([
{
"urn": 100000 + i,
"school_name": f"Brentwood School {i}",
"town": "Brentwood", "local_authority": "Essex",
"postcode": "CM15 8AA", "phase": phase, "year": 202425,
"rwm_expected_pct": 60.0, "attainment_8_score": 50.0,
"ofsted_grade": 2.0, "ofsted_date": None,
}
for i in range(n)
])
def _places_for(monkeypatch, df_factory, urn: int):
from backend import app as app_module
monkeypatch.setattr(app_module, "load_school_data", df_factory)
monkeypatch.setattr(app_module, "get_supplementary_data", lambda db, urn: {})
monkeypatch.setattr(app_module, "_place_registry", None)
client = TestClient(app_module.app, raise_server_exceptions=False)
return client.get(f"/api/schools/{urn}").json()["places"]
def test_a_place_offers_the_phase_page_this_school_appears_on(monkeypatch):
# "primary schools in brentwood" is the query the phase pages exist for,
# and ~950 of them were once reachable by nothing at all.
places = _places_for(monkeypatch, _brentwood_df("Primary"), 100000)
town = next(p for p in places if p["kind"] == "town")
assert town["phases"], "a primary school in a published primary town has a link"
assert town["phases"][0]["url"] == "/schools/brentwood/primary"
assert town["phases"][0]["count"] >= 1
def test_an_all_through_school_offers_both_phase_pages(monkeypatch):
# It genuinely appears on both, so there is no tie to break.
places = _places_for(monkeypatch, _brentwood_df("All-through"), 100000)
town = next(p for p in places if p["kind"] == "town")
assert {p["phase"] for p in town["phases"]} == {"primary", "secondary"}
def test_outcodes_never_offer_a_phase_page(monkeypatch):
# The registry gives outcodes no phase route — nobody searches "primary
# schools in SW11" — and computing them anyway once put a link to a
# nonexistent route on all 1,720 outcode pages.
places = _places_for(monkeypatch, _brentwood_df("Primary"), 100000)
outcode = next((p for p in places if p["kind"] == "outcode"), None)
if outcode is not None:
assert outcode["phases"] == []
def test_a_school_absent_from_the_phase_page_is_not_linked_to_it(monkeypatch):
# The check is URN membership in the registry's own phase list, not a
# re-derivation of the phase mapping. A secondary school must not be sent
# to a primary phase page that does not list it.
from backend.places import MIN_SCHOOLS
def df():
rows = [
{"urn": 100000 + i, "school_name": f"P{i}", "town": "Brentwood",
"local_authority": "Essex", "postcode": "CM15 8AA",
"phase": "Primary", "year": 202425, "rwm_expected_pct": 60.0,
"attainment_8_score": np.nan, "ofsted_grade": 2.0,
"ofsted_date": None}
for i in range(MIN_SCHOOLS)
]
rows.append({
"urn": 900000, "school_name": "Lone Secondary", "town": "Brentwood",
"local_authority": "Essex", "postcode": "CM15 8AA",
"phase": "Secondary", "year": 202425, "rwm_expected_pct": np.nan,
"attainment_8_score": 50.0, "ofsted_grade": 2.0, "ofsted_date": None,
})
return pd.DataFrame(rows)
places = _places_for(monkeypatch, df, 900000)
town = next(p for p in places if p["kind"] == "town")
# The town publishes a primary page, but this secondary school is not on
# it, and there are too few secondaries for a secondary page.
assert town["phases"] == []
def test_the_place_index_rebuilds_when_the_registry_is_replaced(monkeypatch):
"""The reverse index is cached; a stale one would put another dataset's
places on a school page. Invalidation is an identity check against the
registry rather than a second flag, so this asserts the check works."""
from backend import app as app_module
monkeypatch.setattr(app_module, "_place_registry", None)
monkeypatch.setattr(app_module, "_place_index", None)
monkeypatch.setattr(app_module, "_place_index_source", None)
monkeypatch.setattr(app_module, "load_school_data", _brentwood_df("Primary"))
first = app_module.get_place_index()
assert 100000 in first
# Same registry object, so the index is reused rather than rebuilt.
assert app_module.get_place_index() is first
# Drop the registry the way every test that touches place data does. The
# index must follow it, not survive it.
app_module._place_registry = None
monkeypatch.setattr(app_module, "load_school_data",
_brentwood_df("Primary", n=0))
rebuilt = app_module.get_place_index()
assert rebuilt is not first
assert 100000 not in rebuilt, "the index outlived the registry it came from"
-26
View File
@@ -1,26 +0,0 @@
"""
Schema versioning for database migrations.
HOW TO USE:
- Bump SCHEMA_VERSION when making changes to database models
- This triggers an automatic full data reimport on next app startup
WHEN TO BUMP:
- Adding/removing columns in models.py
- Changing column types or constraints
- Modifying CSV column mappings in schemas.py
- Any change that requires fresh data import
"""
# Current schema version - increment when models change
SCHEMA_VERSION = 6
# Changelog for documentation
SCHEMA_CHANGELOG = {
1: "Initial schema with School and SchoolResult tables",
2: "Added pupil absence fields (reading, maths, gps, writing, science)",
3: "Added supplementary data tables: ofsted, parent_view, census, admissions, sen_detail, phonics, deprivation, finance; GIAS columns on schools",
4: "Added Ofsted Report Card columns to ofsted_inspections (new framework from Nov 2025)",
5: "Apply ALTER TABLE additions for RC columns missed by create_all on existing tables",
6: "Removed the Ofsted Parent View feature: dropped fact_parent_view table and model",
}
+29 -126
View File
@@ -1,134 +1,37 @@
# SchoolCompare.co.uk - Project Context
# SchoolCompare project context
## Overview
## Maintained documentation
SchoolCompare is a web application for comparing UK primary school (KS2) performance data. It allows users to:
- Search and browse schools by name, location (postcode), or local authority
- Compare multiple schools side-by-side with charts and tables
- View school rankings by various KS2 metrics
- See historical performance trends across years
Read [README.md](README.md), [docs/ARCHITECTURE.md](docs/ARCHITECTURE.md) and
[docs/DEVELOPMENT.md](docs/DEVELOPMENT.md) for the current implementation.
[docs/LEGACY_CODE.md](docs/LEGACY_CODE.md) records obsolete paths and deliberate
compatibility code. Historical design documents are not current setup instructions.
## Architecture
## Architecture constraints
### Backend (Python/FastAPI)
- **Framework**: FastAPI with uvicorn
- **Database**: PostgreSQL with SQLAlchemy ORM
- **Data Source**: UK Government "Compare School Performance" CSV downloads
Key files:
- `backend/app.py` - Main FastAPI application, API routes
- `backend/config.py` - Configuration via pydantic-settings (env vars, .env file)
- `backend/database.py` - SQLAlchemy engine, session management
- `backend/models.py` - Database models (School, SchoolResult)
- `backend/data_loader.py` - Data queries, geocoding, legacy DataFrame compatibility
- `backend/schemas.py` - Column mappings, metric definitions, LA code mappings
### Frontend (Vanilla JS)
- Single-page application with hash-based routing
- Chart.js for data visualization
- No build step required
Key files:
- `frontend/index.html` - Main HTML structure
- `frontend/app.js` - All application logic, API calls, rendering
- `frontend/styles.css` - Styling (CSS variables, responsive design)
### Database Schema
```
schools school_results
├── id (PK) ├── id (PK)
├── urn (unique, indexed) ├── school_id (FK → schools.id)
├── school_name ├── year (indexed)
├── local_authority ├── rwm_expected_pct
├── school_type ├── reading_expected_pct
├── postcode ├── ... (all KS2 metrics)
├── latitude, longitude └── unique(school_id, year)
└── results → SchoolResult[]
```
## Configuration
Environment variables (or `.env` file):
- `DATABASE_URL` - PostgreSQL connection string (default: `postgresql://schoolcompare:schoolcompare@localhost:5432/schoolcompare`)
- `HOST`, `PORT` - Server binding (default: `0.0.0.0:80`)
- `ALLOWED_ORIGINS` - CORS origins
## Running Locally
1. Start PostgreSQL:
```bash
docker compose up -d db
```
2. Run migration to import CSV data:
```bash
python scripts/migrate_csv_to_db.py --drop
# Add --geocode to geocode postcodes (slower, adds lat/long)
```
3. Start the app:
```bash
uvicorn backend.app:app --host 0.0.0.0 --port 8000
```
## Docker Deployment
```bash
docker compose up -d
```
This starts:
- `db` - PostgreSQL 16 with persistent volume
- `app` - FastAPI application on port 80
## Data
- Source: UK Government Compare School Performance downloads
- Location: `data/` directory with year folders (e.g., `2023-2024/england_ks2final.csv`)
- The `scripts/download_data.py` can fetch data from the government website
## Key Features
- **Location Search**: Enter postcode to find nearby schools (uses postcodes.io API)
- **Multi-school Comparison**: Select multiple schools, view metrics across years
- **Rankings**: Top schools by any KS2 metric, filterable by local authority
- **Variability Analysis**: Shows standard deviation of scores across years
## API Endpoints
- `GET /api/schools` - List/search schools (supports pagination, location search)
- `GET /api/schools/{urn}` - School details with all yearly data
- `GET /api/compare?urns=123,456` - Compare multiple schools
- `GET /api/rankings` - School rankings by metric
- `GET /api/filters` - Available filter options (LAs, types, years)
- `GET /api/metrics` - Metric definitions (single source of truth)
- `GET /api/data-info` - Database stats
- Next.js serves the public UI. FastAPI serves school data from dbt-built `marts.*`.
The backend does not create school tables or import CSVs at startup.
- School coverage spans England and multiple phases, not only primary schools in
Wandsworth and Merton.
- `/api/*` belongs to the FastAPI proxy. Payload uses `/cms-api` and `/admin`.
- Payload runs inside Next.js, with its own `payload` schema and persistent media.
Keep CMS migrations independent of school-data transformations.
- Public and Payload route groups have separate root layouts. Do not introduce
`app/layout.tsx`. Keep site-wide metadata files at the `app/` root.
- Builds must succeed with `DATABASE_URL` unset. Do not call `getCachedPayload()`
at module scope or add DB-backed `generateStaticParams`.
- After changing CMS fields/editors, run `npm run generate:importmap` and commit
the generated import map. See `nextjs-app/docs/PUBLISHING.md`.
- The backend and pipeline GIAS dictionary copies are generated together; preserve
their parity. Tests enforce it.
## SDLC
Full details in `docs/DEPLOY.md`. The short version:
Follow [docs/DEPLOY.md](docs/DEPLOY.md).
- **Never push to `main` directly.** Work on a feature branch and open a PR;
branch protection requires the PR checks (typecheck, tests, builds, AI review)
to pass before merge.
- Merging to `main` deploys automatically **to staging only**: images are
built once, deployed to the staging Portainer stack, and verified by the
Playwright journeys in `e2e/`. Production is a second, manual approval:
the "Promote to Production (manual)" workflow in Gitea Actions, run after
testing the feature on staging. It refuses commits whose staging E2E gate
isn't green. Never trigger it yourself — promotion is the human's call.
- If you change user-facing behaviour, update or extend the `e2e/` journey
tests in the same PR — they gate whether staging is fit for human testing
and whether a commit is promotable.
## Recent Changes
- Added staging environment + automated staging→prod pipeline (Gitea Actions)
- Migrated from CSV file storage to PostgreSQL database
- Added location-based search using postcode geocoding
- Added local authority filter to rankings
- Improved frontend with featured schools, loading states, API caching
# Important
- Do not attempt to start a local server to test the application, it does not work
- Never push directly to `main`. Use a feature branch and a PR with passing checks.
- Merges deploy staging only. Production promotion is a separate human decision;
do not trigger the promotion workflow yourself.
- Update E2E journeys in the same PR when changing user-facing behaviour.
- Do not attempt to start a local server to test the application; use unit checks
and the configured integration environment.
+16
View File
@@ -18,6 +18,10 @@
# TYPESENSE_SEARCH_KEY — Typesense search-only key (exposed to frontend)
# UNLEASH_URL — http://<unleash-ip>:4242/api (empty = all flags off)
# UNLEASH_API_TOKEN — Unleash *client* token, environment: development
# PAYLOAD_SECRET — Payload CMS encryption secret. REQUIRED: long and
# random, and DIFFERENT from production's. Sharing
# it would let a staging session authenticate
# against production.
# AIRFLOW_ADMIN_USER — Airflow admin username (default: admin)
# AIRFLOW_ADMIN_PASSWORD — Airflow admin password. REQUIRED: the api-server
# refuses to start without it, rather than falling
@@ -89,9 +93,20 @@ services:
- FASTAPI_URL=http://backend:80/api
- TYPESENSE_URL=http://typesense:8108
- TYPESENSE_API_KEY=${TYPESENSE_SEARCH_KEY:-changeme}
# Payload CMS runs inside this container, in the `payload` schema of the
# staging database. Staging has its own stack, its own Postgres and its
# own admin account — never production's.
- DATABASE_URL=postgresql://${DB_USERNAME}:${DB_PASSWORD}@sc_database:5432/${DB_DATABASE_NAME}
- PAYLOAD_SECRET=${PAYLOAD_SECRET:?set PAYLOAD_SECRET in the staging Portainer stack environment}
volumes:
# Portainer prefixes volume names with the stack name, so this is
# automatically isolated from production's media.
- payload_media:/app/media
depends_on:
backend:
condition: service_healthy
sc_database:
condition: service_healthy
networks:
backend: {}
macvlan:
@@ -242,3 +257,4 @@ volumes:
typesense_data:
airflow_logs:
unleash_cache:
payload_media:
+16
View File
@@ -9,6 +9,9 @@
# TYPESENSE_SEARCH_KEY — Typesense search-only key (exposed to frontend)
# UNLEASH_URL — http://<unleash-ip>:4242/api (empty = all flags off)
# UNLEASH_API_TOKEN — Unleash *client* token, environment: production
# PAYLOAD_SECRET — Payload CMS encryption secret. REQUIRED: long and
# random. Changing it invalidates every admin
# session. Staging MUST use a different value.
# AIRFLOW_ADMIN_USER — Airflow admin username (default: admin)
# AIRFLOW_ADMIN_PASSWORD — Airflow admin password. REQUIRED: the api-server
# refuses to start without it, rather than falling
@@ -78,9 +81,21 @@ services:
- FASTAPI_URL=http://backend:80/api
- TYPESENSE_URL=http://typesense:8108
- TYPESENSE_API_KEY=${TYPESENSE_SEARCH_KEY:-changeme}
# Payload CMS runs inside this container. It reaches Postgres over the
# `backend` network and keeps its tables in the `payload` schema, so no
# pipeline operation on `public` can touch blog content.
- DATABASE_URL=postgresql://${DB_USERNAME}:${DB_PASSWORD}@sc_database:5432/${DB_DATABASE_NAME}
# Same :? form as AIRFLOW_ADMIN_PASSWORD: refuse to start rather than
# boot with an empty secret and silently accept forged sessions.
- PAYLOAD_SECRET=${PAYLOAD_SECRET:?set PAYLOAD_SECRET in the Portainer stack environment}
volumes:
# Blog images. Not reproducible from the pipeline — must be backed up.
- payload_media:/app/media
depends_on:
backend:
condition: service_healthy
sc_database:
condition: service_healthy
networks:
backend: {}
macvlan:
@@ -231,3 +246,4 @@ volumes:
typesense_data:
airflow_logs:
unleash_cache:
payload_media:
+107
View File
@@ -0,0 +1,107 @@
# Architecture
This describes the implementation as reviewed on 2026-09-14. It distinguishes
current behaviour from improvements still to be implemented.
## Request flow
```text
Browser → Next.js public routes
├─ /api/* proxy → FastAPI → cached DataFrames / PostgreSQL marts
│ ├─ Typesense (search and suggestions)
│ └─ postcodes.io (postcode lookup)
└─ /admin, /cms-api, /blog → Payload → payload schema + media volume
Next.js server rendering → FastAPI directly through FASTAPI_URL
```
`nextjs-app/lib/api.ts` contains typed fetch wrappers and revalidation defaults.
The proxy is `nextjs-app/app/(frontend)/api/[...path]/route.ts`. Payload uses
`/cms-api` so its routes do not collide with the FastAPI proxy. The proxy denies
`/api/flags`; server-side rendering reads flags directly from FastAPI.
## Data ownership
| Layer | Owner and role |
|---|---|
| Source data | GIAS, DfE EES, Ofsted, finance, deprivation and council admission-distance sources |
| `raw` | Singer taps and the PostgreSQL target configured in `pipeline/meltano.yml` |
| Staging/intermediate/marts | dbt models in `pipeline/transform`; marts are materialized tables |
| `marts.dim_school`, `marts.dim_location` | School identity and location, filtered to supported England establishments |
| `marts.fact_*` | Performance and supplementary datasets; coverage and years vary |
| Typesense `schools` alias | Search documents built by `pipeline/scripts/sync_typesense.py` |
| `payload` | CMS collections and migrations in `nextjs-app/`; independent of dbt |
| Media volume | Uploaded blog media; requires backup and cannot be regenerated from school datasets |
`backend/models.py` maps existing marts for reading. It does not create the school
schema. There is no startup schema-version migration or CSV reimport. Payload's
`nextjs-app/migrations/` is active and must not be confused with the removed
legacy backend migration code.
Coordinates normally come from GIAS British National Grid coordinates transformed
by PostGIS in `dim_location.sql`. `pipeline/scripts/geocode_postcodes.py` is a
manual fallback utility, not a task wired into the current school-data DAG.
Backend postcode searches also use postcodes.io; that lookup does not populate
school coordinates in the database.
## Backend boundaries
- `app.py`: routes, middleware, search filtering, sitemap/place publication and response assembly.
- `data_loader.py`: SQL loading, process-local DataFrame caches, Typesense calls,
postcode lookups, supplementary queries and benchmark calculation.
- `database.py`: synchronous SQLAlchemy engine and sessions.
- `schemas.py`: metric definitions, column mappings and display metadata; despite
its name this is not a collection of Pydantic API response models.
- `places.py` and `localities.py`: place registry and curated locality information.
- `flags.py`: Unleash-backed feature flags, disabled when no server is configured.
- `gias_codes.py` / `ofsted_codes.py`: source-code translation and display rules.
Search starts from a cached latest-row-per-school snapshot. Detail pages read
history from the full DataFrame and supplementary data from marts. Comparisons
batch supplementary queries across selected URNs. Async routes still contain
synchronous dependency calls; a fully asynchronous database layer is not present.
## Frontend boundaries
`app/(frontend)` owns the public root layout and pages. `app/(payload)` owns the
CMS root layout. Do not add a shared `app/layout.tsx`: these groups deliberately
have separate root layouts. Root metadata files remain in `app/`.
Server pages fetch initial data and pass it to client views. Client state uses
React hooks, URL search parameters and the comparison context/localStorage.
There is no SWR dependency. Leaflet maps are loaded through dynamic wrappers;
Chart.js renders performance and comparison charts.
`components/school/` contains detail sections, with section decisions and data
preparation in `lib/schoolSections.ts`. `lib/types.ts` contains manually maintained
API types. `payload-types.ts` and the Payload import map are generated artifacts.
## Publication and caching today
1. Airflow DAGs extract and validate source data, then run selected dbt builds.
2. Relevant DAGs rebuild Typesense and swap the `schools` alias.
3. They call `POST /api/admin/reload` with `X-API-Key` to refresh school DataFrames.
4. A separate weekly sitemap DAG calls `POST /api/admin/regenerate-sitemap`,
rebuilding places and sitemaps.
GIAS is scheduled daily, Ofsted monthly, and annual datasets are manually
triggered. The DAG definitions are authoritative for selectors and dependencies.
Caches exist in several independent layers: backend DataFrames and registries,
backend HTTP Cache-Control/ETags, Next.js fetch/page revalidation, and browser or
shared HTTP caches where configured. Place fetches request a one-week revalidation
interval. HTTP ETags are computed after route execution, not before database work.
Known limitations: reload clears the old DataFrames before verifying replacement
data; places/sitemaps refresh separately; Next.js caches are not explicitly purged
by the pipeline; Typesense import results are not validated before alias publication.
Do not describe this sequence as an atomic dataset release. These are follow-up
reliability tasks, not changes implemented by the documentation cleanup.
## Deployment references
See [DEPLOY.md](DEPLOY.md). PR checks include frontend typechecking/tests, backend
unit tests, image builds and AI review. Staging journeys run after merging.
Production promotion retags a selected commit's images. Current health polling
checks HTTP success, not the deployed commit identity; overlapping staging runs
remain a release-verification concern.
+94
View File
@@ -0,0 +1,94 @@
# Development and validation
## Prerequisites and environment boundaries
Use a feature branch. The deployed stack is the integration environment; do not
assume a local server can run from a fresh checkout. This cleanup did not start
local servers or provision databases. Unit tests use fixtures and mocks.
The current versions are not yet aligned:
| Component | Container | PR checks |
|---|---|---|
| Backend | Python 3.11 | Python 3.12 |
| Frontend | Node 24 | Node 22 |
| Pipeline | Python 3.13 | Pipeline image build |
Use the component's container version when reproducing deployment behaviour.
The backend dependency pins predate Python 3.14; do not assume the system Python
can install or run them. Version alignment is a separate maintenance task.
## Frontend checks
```sh
cd nextjs-app
npm ci
npm run typecheck
npm test -- --runInBand
```
`npm run build` is the production build check. There is no `lint` script.
Tests live in `__tests__/` and use Jest/React Testing Library. These checks do not
prove that live PostgreSQL queries, Typesense or a deployed proxy work.
The frontend `.env.example` documents runtime variables. Browser traffic normally
uses `/api`; `FASTAPI_URL` is an absolute server-side URL ending in `/api`.
Payload additionally needs `DATABASE_URL` and `PAYLOAD_SECRET` when used at runtime.
Never commit credentials or real `.env` files.
## Backend checks
From the repository root, using an available Python 3.11 or 3.12 interpreter:
```sh
python3.11 -m venv /tmp/schoolcompare-backend-venv
/tmp/schoolcompare-backend-venv/bin/python -m pip install -r requirements.txt pytest 'httpx<0.28'
/tmp/schoolcompare-backend-venv/bin/python -m pytest backend/tests -q
```
Substitute `python3.12` if matching PR CI. The test dependencies above match the
current workflow; they are not yet captured in a dedicated development lockfile.
Backend configuration is defined in `backend/config.py`; `.env.example` documents
commonly used values. `ALLOWED_ORIGINS` uses a JSON array, not a comma-separated string.
## Data and pipeline work
The app needs populated `marts.*` tables. A new Postgres instance alone is not a
working school-data environment. Use the existing managed pipeline or an approved
snapshot; the removed CSV importer cannot build the current schema.
The pipeline container includes Meltano, dbt/Postgres, Airflow and the custom taps.
Airflow commands/selectors live in `pipeline/dags/`. Schema tests live in
`pipeline/transform/tests/` and model YAML files. Run the relevant `dbt build`
selector in an isolated data environment for model changes; it writes tables and
is not a read-only smoke test. Prefer `python -m dbt.cli.main` as the DAGs do.
GIAS dictionaries are generated together by
`pipeline/scripts/generate_gias_codes.py`. The backend and pipeline copies are
intentional; `backend/tests/test_gias_codes.py` checks that they stay identical.
For Payload collection/editor changes, run `npm run generate:importmap` in
`nextjs-app/` and include the generated map. Preserve CMS migrations and the
separate `payload` schema. See [publishing](../nextjs-app/docs/PUBLISHING.md).
## End-to-end checks
Against an existing, authorised test environment:
```sh
cd e2e
npm ci
npx playwright install chromium
BASE_URL=https://your-test-environment.example npx playwright test
```
The suite does not start a web server. CI installs Chromium with system dependencies
and runs against staging. Use the configured staging target: `docs/DEPLOY.md`
records the public staging proxy limitation. User-visible behaviour changes should
update the corresponding journeys.
## Before requesting review
Run checks relevant to the change, inspect `git diff --check`, and report checks
that could not run. Do not publish or promote as part of local validation.
[DEPLOY.md](DEPLOY.md) documents the PR and human promotion gates.
+89
View File
@@ -0,0 +1,89 @@
# Legacy and unused-code inventory
Reviewed 2026-09-14. This inventory records source evidence, not production usage
telemetry. A command with no repository caller may still be run manually or from
an external scheduler. Historical specs and prototypes are not runtime imports.
## Method and scope
Searched backend imports, tests, CLI scripts, Airflow DAGs, Meltano configuration,
Gitea workflows, Dockerfiles and documentation. For frontend candidates, inspected
TypeScript imports, re-exports, literal dynamic imports and `require` calls,
resolving relative and `@/` paths while excluding tests, dependencies and build
output. Checked candidates again with text searches including tests.
Next.js route files, generated Payload import-map entries and plugin discovery
are entry points even without ordinary imports. This is why a zero-import count
alone is not sufficient grounds for deletion. Computed imports and external
operators are outside this static audit.
## Removed in this cleanup
These names are recorded for Git-history lookup; they are no longer file links.
| Removed path or symbol | Evidence and replacement |
|---|---|
| `backend/migration.py` | Imported `School` and `SchoolResult`, which no longer exist in `backend/models.py`. Only the legacy CSV CLI imported it. Current tables are built by dbt. |
| `backend/version.py` | Only the legacy importer consumed `SCHEMA_VERSION`. FastAPI lifespan does not perform version-triggered imports. This is unrelated to active Payload migrations. |
| `scripts/migrate_csv_to_db.py` | Imported removed `init_db`/`set_db_schema_version` helpers and the obsolete models indirectly. No runtime, DAG or workflow calls it. Use the managed pipeline for current marts. |
| `scripts/geocode_schools.py` | Imported the removed `School` ORM model. No pipeline/workflow calls it. Coordinates now come from GIAS/PostGIS; a separate mart-aware manual utility remains under `pipeline/scripts/`. |
| `backend.data_loader.haversine_distance` | No callers. Search uses its inline vectorised NumPy calculation. |
| `nextjs-app/lib/api.ts: fetcher` | No callers; SWR is not installed. Application fetches use the named API wrappers. |
| `nextjs-app/lib/api.ts: kmToMiles` | No callers. `calculateDistance` remains because `CutoffMapPanel` uses it. |
The removed command files could not import successfully against the current
backend. This cleanup does not run replacements, migrate data or modify databases.
Their previous implementations remain recoverable from Git history.
## Unused candidates retained for a separate cleanup
| Candidate | Evidence | Recommended next step |
|---|---|---|
| `nextjs-app/components/LoadingSkeleton.tsx` and its CSS | No application or test imports found. | Remove together after confirming no planned use. |
| `nextjs-app/components/Pagination.tsx` and its CSS | No application or test imports found; HomeView implements load-more behaviour. | Remove as a pair if numbered pagination will not return. |
| `nextjs-app/components/SchoolCard.tsx` and its CSS | Imported by its own tests, not application code. HomeView uses SchoolRow/SecondarySchoolRow. | Decide whether to retire the card design; if removed, remove its dedicated tests as well. Passing tests do not establish runtime use. |
| `backend/database.py: get_db`, `get_db_session` | No remaining callers after removing the importer. Current code creates SessionLocal directly. | Either adopt these helpers during session-lifecycle cleanup or remove them; do not rewrite active sessions in a documentation change. |
| `backend/schemas.py: COLUMN_MAPPINGS`, `NULL_VALUES`, `LA_CODE_TO_NAME` | No remaining Python consumers found after importer removal. Other constants in this module are active. | Remove individual constants after checking external data utilities; retain the module. |
| `backend/config.py: data_dir`, `max_page_size`, `rate_limit_burst` | No active consumers found. `default_page_size` appears only in a branch that expects None, although the route supplies a concrete default. | Reconcile settings with route validation in a focused API change. |
## Legacy/manual paths requiring operational verification
| Path | Status and reason to retain for now |
|---|---|
| FastAPI `/`, `/compare`, `/rankings`, `/favicon.svg`, `/robots.txt`, and conditional `/static` | Old frontend-serving routes reference a `frontend/` directory absent from the checkout and backend image. Next.js owns these public surfaces. Removal changes externally callable routes, so first check proxy/operator usage and define replacement responses. |
| `scripts/fetch_real_data.py`, `scripts/download_data.py` | Historical standalone CSV utilities. The fetch script targets Wandsworth/Merton; neither is wired into the managed pipeline. Marked historical, retained pending confirmation of manual use. |
| `pipeline/scripts/geocode_postcodes.py` | Mart-aware postcode fallback, not called by the current DAGs. Do not confuse it with the removed legacy ORM geocoder. Verify the target schema before manual use. |
| `docker-compose.yml` | Uses unpublished `:latest` release tags and lacks frontend Payload DB/secret/media configuration. Retained as an old development topology, not recommended onboarding. |
| `nextjs-app/docker-compose.yml` | Standalone legacy recipe with old backend port assumptions and no CMS persistence setup. Retained until its consumers are checked. |
| `MIGRATION_SUMMARY.md`, `docs/superpowers/`, `mockups/` | Historical designs and prototypes. Retain as history; do not follow as current deployment instructions. |
| `scripts/sql/drop_fact_parent_view.sql` | One-off maintenance SQL. Not an application entry point; repository call-site searches cannot establish whether it is still needed operationally. |
## Active code that can look obsolete
- `backend/data_loader.py` older-mart query fallbacks are covered by backend tests
and support databases at different migration stages. Remove only after verifying
the deployed schemas in every supported environment.
- `backend/gias_codes.py` and `pipeline/scripts/gias_codes.py` are intentionally
generated copies for separate runtime images. Their parity is tested.
- `nextjs-app/migrations/`, `payload-types.ts` and the Payload import map are active
CMS artifacts, not remnants of the removed school importer.
- `get_available_years`, `get_available_local_authorities` and `get_schools_count`
in `data_loader.py` are called through `get_data_info`, which serves the backend
data-info endpoint. They are not dead functions.
- `get_supplementary_data` is an intentional single-school wrapper around the
batch implementation.
- `pipeline/transform` models named `legacy` can be active data sources: annual
DAG selectors explicitly include legacy KS2/KS4 lineage. Names alone do not
establish obsolescence.
## Suggested next passes
1. Decide the fate of the three unused UI components and remove paired assets/tests.
2. Consolidate backend session usage and remove abandoned settings/constants.
3. Verify external consumers, then retire static-serving API routes and old compose recipes.
4. Audit manual data utilities with pipeline operators before deleting them.
5. Revisit compatibility fallbacks only after documenting supported schema versions.
Validation for this cleanup should include frontend typechecking/tests, Python
syntax checks, reference searches and documentation link checks. Live database,
external scheduler and deployed route usage require separate integration evidence.
File diff suppressed because it is too large. Load diff
@@ -0,0 +1,368 @@
# Giving schoolcompare a human author: an About page and a blog
**Date:** 2026-09-02
**Status:** Design — awaiting review
**Scope:** A named author for the site, an `/about` page, and a Payload-CMS-backed
blog at `/blog`.
## Why
The site reads as synthetic. Not because of its tone, but because of three
specific absences:
1. **Nobody is accountable for the numbers.** There is no author, no statement
of why the site exists, and no one who can be wrong. The only human trace on
the entire site is `contact@schoolcompare.co.uk` in the footer.
2. **No visible judgement.** Every figure is presented as though it fell out of
a machine. Hundreds of editorial decisions went into this codebase — which
metrics to show, when a benchmark is invalid, what to suppress — and not one
of them is visible to a reader. `isSpecialSchool()` silently drops the
England comparison for special schools and PRUs because that comparison is
meaningless; nowhere does the site *say* so.
3. **The voice is institutional third person.** "schoolcompare brings it all
into one place." "Built for parents, governors, journalists." That is
brochure register, and it is precisely the register that machine-generated
content defaults to.
There is a second, independent reason. The SEO programme
(`2026-08-20-seo-programme-design.md`) defines eight workstreams and none of
them address E-E-A-T or authorship. School performance data is YMYL territory;
an anonymous site republishing DfE figures has no authorship signal at all. This
work fills that hole, and the blog gives W6 (explainer content) somewhere to
live.
### The failure mode to avoid
The standard fix — a stock photo and "Hi, I'm Tudor, and I'm passionate about
education!" — reads as *more* synthetic than the current coldness. Manufactured
warmth is a stronger machine-tell than plain institutional voice. Everything
here has to be specific, occasionally awkward, and willing to be unflattering,
or it makes the problem worse.
## Positioning
The author is **Tudor**: first name only, real photograph, no surname, no
employer named.
The credibility claim is deliberately **not** educational expertise. The About
page states plainly: *"I'm not an education expert."* Authority comes from two
things that are actually true:
- **Experience.** A parent going through primary admissions in south-west London
right now. Google's E-E-A-T leads with Experience, and lived experience of the
thing is exactly what the DfE's own service lacks.
- **Method.** Every number's provenance is stated, so a reader can check the
site rather than trust it.
This is more durable than borrowed expertise: it cannot be undermined by someone
noticing the author has no teaching qualification.
**Consequence for the design.** A `Person` entity with no surname is a weak
search signal and cannot be corroborated off-site. The credibility load
therefore shifts onto the methodology being visibly rigorous. That is a design
constraint, not a caveat — it is why the About page carries a substantial
"how this is built and where it can be wrong" section rather than a short bio.
### Voice rules
Applied to About and every post. Recorded here so the voice does not drift.
- First person singular. "I built", not "we provide".
- Concrete over general. "when we were looking at schools in Wandsworth" beats
any amount of stated warmth.
- State limits before someone else finds them. Every post that presents a
metric says what it does not show.
- No mission statements, no "passionate about", no invented team.
- No em dashes. One of the clearest tells of machine-written prose, which is
the exact problem this work exists to fix.
- Short sentences. The existing code comments in this repo are already written
this way; the prose should match.
## Scope
**In:**
- `/about` — a coded page (not CMS-managed).
- `/blog` and `/blog/[slug]` — Payload-backed, with an index and post pages.
- Payload CMS installed into the existing Next application.
- Footer and navigation links to both.
- `Person`, `Organization`, `BlogPosting`, `BreadcrumbList` JSON-LD.
- RSS feed and sitemap integration.
- One first post, so the blog does not launch empty.
**Out (deliberately):**
- Rewriting existing homepage/how-it-works copy into first person. Worth doing,
but it would double the review surface of this PR. Separate change.
- In-product signed notes on school pages (the "distributed humanity" idea).
Revisit once About and the blog exist.
- Comments, newsletter, author accounts beyond one.
- A team page. There is no team.
## Architecture
### Topology
Payload 3 installs **into the existing Next application** and serves `/admin`
from the same container. One image, one deploy, no new service. This is
Payload 3's native model and it makes on-demand revalidation trivial, because
the CMS hooks run in the same process as the Next cache.
Accepted costs: the public site's image now carries Payload, so a CMS security
patch redeploys the whole site; and the image grows substantially.
### Two collisions that must be handled
**1. `/api` is already taken.** `app/api/[...path]/route.ts` is a catch-all that
proxies `/api/*` to FastAPI at runtime. Payload's default API route is also
`/api`. Left alone, these fight, and the failure is not clean — the catch-all
would swallow Payload's admin API calls and forward them to FastAPI.
Payload's API route is therefore remapped:
```ts
routes: { api: '/cms-api', admin: '/admin' }
```
with its route group at `app/(payload)/cms-api/[...slug]/route.ts`. The
`/cms-api` prefix must also be added to the FastAPI proxy's excluded-paths list
as a defensive second line.
**2. `next.config.js` is CommonJS.** Payload's `withPayload()` wrapper is ESM
only. The config must become `next.config.mjs`, converting `module.exports` to
`export default` and wrapping the export. All existing content — the standalone
output, `outputFileTracingIncludes`, the staging `X-Robots-Tag` header block,
the CSP — carries over unchanged. This is mechanical but it touches the file
that controls staging's noindex, so it needs care and an explicit test.
### Database
Payload uses the existing `sc_database` Postgres instance, in its **own
`payload` schema**:
```ts
db: postgresAdapter({
pool: { connectionString: process.env.DATABASE_URL },
schemaName: 'payload',
})
```
The frontend container is already on the `backend` Docker network, so it can
reach `sc_database:5432` with no networking change. It needs a new
`DATABASE_URL` environment variable.
Schema isolation is not cosmetic. `public` currently holds the application
tables and Airflow's metadata, and `scripts/migrate_csv_to_db.py --drop` exists
to drop and reimport. Blog content living in its own schema means no data
pipeline operation can destroy it.
**Verified 2026-09-02** (this was an open question when the spec was written).
`--drop` calls `run_full_migration()` in `backend/migration.py`, which drops
exactly two tables by name:
```python
ks2_tables = ["school_results", "schools"]
for tname in ks2_tables:
if tname in existing:
Base.metadata.tables[tname].drop(bind=engine)
```
There is no `Base.metadata.drop_all()` anywhere in `backend/`, and no
`DROP SCHEMA`. The only other drop is `_apply_schema_drops()`, a single
schema-qualified `DROP TABLE IF EXISTS marts.fact_parent_view CASCADE`.
Nothing sets `search_path`, so the SQLAlchemy metadata resolves to `public`,
and `inspector.get_table_names()` does not even enumerate other schemas.
So the guarantee is stronger than schema isolation alone: `--drop` targets two
named tables that Payload does not have, and would not reach `posts`, `media`
or `users` even if they shared a schema. The `payload` schema remains the right
choice — it protects against a *future* broadening of that script rather than
today's behaviour — but the safety claim rests on verified code, not on
assumption.
Putting CMS tables in this instance is consistent with existing practice —
Airflow already stores its metadata there.
### Migrations
Payload's Postgres adapter auto-pushes schema in development and requires
explicit migrations in production. Use `prodMigrations`, which runs pending
migrations during server initialisation:
```ts
db: postgresAdapter({ /* ... */, prodMigrations: migrations })
```
This is preferred over a one-shot init container (the `airflow-init` pattern)
because the app is a single long-running process and there is no ordering
problem to solve. Migration files are generated with `payload migrate:create`
and committed, so schema changes travel through the same PR and staging gate as
code.
### Media
Uploads go to a Docker named volume, consistent with `postgres_data`,
`typesense_data` and `airflow_logs`.
- `staticDir` must be an **absolute** path in Payload 3: `/app/media`.
- The container runs as `nextjs` (uid 1001). The Dockerfile must
`mkdir -p /app/media && chown nextjs:nodejs /app/media` **before** the volume
is mounted, or Docker will create the mountpoint root-owned and every upload
will fail with EACCES.
- `sharp` moves from `devDependencies` to `dependencies` — Payload needs it at
runtime to generate `imageSizes`.
- The volume must be added to the backup routine alongside Postgres. A blog
post's images are not reproducible from the pipeline.
### Rendering
**Constraint:** CI builds the image with no database reachable. Blog pages
therefore cannot use build-time `generateStaticParams` — that would either fail
the build or bake in an empty post list.
Instead: ISR. Post and index pages declare a `revalidate` window and render on
first request, with Payload `afterChange` / `afterDelete` hooks calling
`revalidatePath('/blog')` and `revalidatePath('/blog/' + slug)` for immediate
publication. Because Payload runs in the same process, the hook calls
`revalidatePath` from `next/cache` directly — no webhook, no shared secret.
The ISR cache lives on container disk and is cleared by a redeploy. For a
single container serving a handful of posts this is fine.
### Collections
- **`posts`** — `title`, `slug`, `publishedAt`, `excerpt`, `heroImage`
(relation to `media`), `content` (Lexical rich text), `seo` group
(`metaTitle`, `metaDescription`), `_status` (drafts enabled).
- **`media`** — upload collection, `alt` required, `imageSizes` for thumbnail
and hero widths, public read access.
- **`users`** — Payload's auth collection. One account. Public creation
disabled.
Drafts are enabled so posts can be written over several sittings and previewed
before publication.
**Payload Blocks** are how posts embed live product components — a real trend
chart or comparison table inside a post, rendered from live data rather than
screenshotted. This is the main thing the CMS has to earn back against
file-based MDX, and it directly serves the goal: showing judgement in context.
Ship with one block (a callout/aside for "what this number doesn't tell you");
add a live-chart block once a post needs it.
### Security
`/admin` is the first authenticated surface on this site. Public, hardened:
- `PAYLOAD_SECRET` — long, random, set in the Portainer stack environment, never
committed. The same variable must exist in staging with a *different* value.
- Strong unique password on the single admin account.
- Login rate limiting via Payload's `maxLoginAttempts` / `lockTime`.
- `X-Robots-Tag: noindex, nofollow` on `/admin/*` and `/cms-api/*`, and a
`robots.ts` disallow. The admin panel must never be indexed.
- Public user creation disabled; no open registration.
- Verify the existing CSP `frame-ancestors` directive does not break the admin
panel.
Residual risk, accepted: a future Payload authentication CVE is live against the
public internet. Mitigation is prompt patching, which the staging→prod pipeline
already supports. If this becomes uncomfortable, restricting `/admin` at the
proxy to LAN/VPN is a one-line change later.
Staging note: staging runs the same image on `stx.`, so it gets its own admin
panel and its own database. It must have its own `PAYLOAD_SECRET` and its own
credentials — never production's.
## Deployment changes
- `nextjs-app/Dockerfile` — create and chown `/app/media`; ensure Payload's
admin bundle and `sharp` survive standalone output file tracing.
- `docker-compose.portainer.yml` and the staging equivalent — add
`DATABASE_URL` and `PAYLOAD_SECRET` to the `frontend` service, add a
`payload_media` volume mounted at `/app/media`, and add
`depends_on: sc_database`.
- Document both new environment variables in the compose header comment block,
which is where this stack records its configuration.
## SEO
- `Person` (Tudor, with photo) and `Organization` JSON-LD on `/about`.
- `BlogPosting` + `BreadcrumbList` on post pages, with `author` referencing the
same `Person`.
- Canonical URLs on `/blog` and every post.
- Posts and `/about` added to the existing sitemap (`app/sitemap.xml/route.ts`
and `app/sitemaps/[...parts]`). Post URLs come from Payload at request time.
- RSS feed at `/blog/rss.xml`.
- Footer links to both pages, under a new "About" column.
**Navigation is deliberately left alone.** `Navigation.tsx` renders a bottom tab
bar on mobile that already carries four items (Search, Compare, Rankings,
Admissions). A fifth tab makes each one cramped at 320px, and About and Blog are
both lower-intent than any of the four. Both live in the footer; About
additionally gets a byline link from every post, which is where a reader who
cares actually asks the question. Revisit only if analytics show people hunting
for it.
## Testing
Unit (Jest):
- Post rendering, including a post with no hero image and one with no excerpt.
- Slug generation and collision handling.
- JSON-LD shape for `BlogPosting` and `Person`.
- The `next.config.mjs` conversion preserves the staging `X-Robots-Tag` rule —
this guards the riskiest mechanical change in the plan.
E2E (Playwright, `e2e/`, required by CLAUDE.md for user-facing change):
- `/about` renders, shows the author name and photo, and is reachable from the
footer and nav.
- `/blog` lists at least one post; clicking through reaches the post.
- A post page renders title, date, body and byline.
- `/admin` responds with `noindex` and does not leak a stack trace when
unauthenticated.
Note the known constraint: new journeys cannot be proven in PR checks, because
the staging E2E gate runs post-merge.
## Risks
| Risk | Mitigation |
|---|---|
| `next.config.mjs` conversion silently drops the staging noindex header, making staging a crawlable duplicate | Unit test asserting the header rule; verify on staging before promotion |
| Payload API route collides with the FastAPI `/api` proxy | Remap to `/cms-api`; add to the proxy's exclusion list |
| Media volume mounts root-owned; all uploads fail with EACCES | `mkdir`+`chown` in the Dockerfile before the mount; test an upload on staging |
| Build fails or bakes empty content because CI has no DB | No build-time DB access; ISR only |
| A pipeline `--drop` destroys blog content | Separate `payload` schema; verify `--drop` blast radius before building |
| Media volume not backed up; images unrecoverable | Add `payload_media` to the backup routine |
| Payload auth CVE exposed publicly | Prompt patching; proxy restriction available as a fallback |
| Blog launches empty or goes stale | Ship with one post; cadence is explicitly "a few times a year", so no cadence is promised anywhere on the page — no dates implying a schedule |
## Sequence
Each step is independently reviewable and mergeable.
1. **Payload foundation** — install, `next.config.mjs` conversion, `payload`
schema, `/cms-api` remap, `users` collection, `/admin` hardening, compose and
Dockerfile changes. No public-facing change yet. Verify on staging that the
site is unchanged and `/admin` works.
2. **`/about`** — coded page, photo, `Person`/`Organization` JSON-LD, footer and
nav links, e2e journey. Independently valuable and does not depend on the
blog.
3. **Blog** — `posts` and `media` collections, `/blog` index and post pages, ISR
plus revalidation hooks, RSS, sitemap, structured data, e2e journeys.
4. **First post** — written in the admin panel, published through the normal
flow, proving the whole path end to end.
Step 1 carries all the infrastructure risk and none of the visible benefit, so
it should be verified on staging carefully before step 2 starts.
## Dependencies on Tudor
- **A photograph.** Blocks step 2. Nothing else in the plan is blocked by it.
- **The first post's subject.** Blocks step 4 only. Suggested: what school
performance data cannot tell you — it demonstrates judgement, is genuinely
useful, and is the kind of thing an anonymous or machine-written site will not
publish.
- ~~Confirmation that `scripts/migrate_csv_to_db.py --drop` is schema-scoped.~~
**Resolved 2026-09-02** — verified in `backend/migration.py`; see the
Database section. No action needed.
+180
View File
@@ -1935,6 +1935,63 @@ async function firstPlaceOfKind(page: Page, kind: string) {
return hit as { kind: string; slug: string; name: string; count: number };
}
/**
* The round trip. Place pages always linked down to school pages; school
* pages linked nowhere on the site, so the ~27k of them that carry most of
* the inbound authority stranded it — their only anchor pointed at the
* school's own website.
*
* Asserting both directions is the point. A one-way link is what already
* existed and is not what this journey is for.
*/
test('a school page links back into the location layer, and the place page links down', async ({ page }) => {
const town = await firstPlaceOfKind(page, 'town');
// Start from the place page and take its first school, so the pair is
// guaranteed to be genuinely related rather than a hardcoded guess.
await page.goto(`/schools/${town.slug}`);
const schoolHref = await page.locator('a[href^="/school/"]').first()
.getAttribute('href');
expect(schoolHref, 'the town page listed no school to follow').toBeTruthy();
await page.goto(schoolHref!);
// Down: the school page must offer a link back to the town it sits in.
const backToTown = page.locator(`a[href="/schools/${town.slug}"]`);
await expect(backToTown).toHaveCount(1);
await expect(backToTown).toBeVisible();
// The anchor says what it leads to, which is worth more than "see more".
await expect(backToTown).toContainText(town.name, { ignoreCase: true });
await expect(backToTown).toContainText(/\d+ schools?/);
// And the breadcrumb resolves the school into a real hierarchy.
const blocks = await page.locator('script[type="application/ld+json"]')
.allTextContents();
const graph = blocks.join(' ');
expect(graph).toContain('"BreadcrumbList"');
// The narrower type, not the EducationalOrganization parent it used to be.
expect(graph).toContain('"School"');
/*
* The phase variants are the pages this most needs to reach: ~950 of them
* were once reachable by nothing at all, absent from every sitemap and
* unlinked from the place page. Conditional because not every school sits
* in a town that publishes one.
*/
const phaseLink = page.locator(`a[href^="/schools/${town.slug}/"]`).first();
if (await phaseLink.count()) {
const phaseHref = await phaseLink.getAttribute('href');
expect((await page.request.get(phaseHref!)).status()).toBe(200);
await expect(phaseLink).toContainText(/primary|secondary/);
}
// Following it lands on a real page, not a 404.
await backToTown.click();
await page.waitForURL(new RegExp(`/schools/${town.slug}$`));
await expect(page.locator('h1')).toContainText(town.name, { ignoreCase: true });
});
for (const [kind, prefix, article] of [
['town', '/schools/', 'a'],
['authority', '/schools/authority/', 'an'],
@@ -2554,3 +2611,126 @@ test('the destinations section never claims a pupil stayed at this school', asyn
const text = (await section.textContent()) ?? '';
expect(text).not.toMatch(/stayed on (here|at this school)/i);
});
/**
* The About page and the blog exist to give the site a named human author.
* These journeys assert the load-bearing parts of that — a name, a face, the
* honesty claim, and a resolvable Person entity — rather than exact copy,
* which will be edited.
*
* Both are behind flags (about_page, blog), so each has a lit journey and a
* dark one. Flag state is read from the observable effect rather than from
* /api/flags, which the public proxy denies on purpose — the same approach
* distanceFeatureIsOn() takes above.
*/
async function aboutPageIsOn(page: Page): Promise<boolean> {
return (await page.request.get('/about')).ok();
}
async function blogIsOn(page: Page): Promise<boolean> {
return (await page.request.get('/blog')).ok();
}
test('with the about page off, it is absent rather than empty', async ({ page }) => {
test.skip(await aboutPageIsOn(page), 'the about_page flag is on in this environment');
// Dark means the URL does not exist, not that it renders empty: a 404 is
// what stops a crawler keeping the page in its index.
expect((await page.request.get('/about')).status()).toBe(404);
// A footer link into a 404 is the failure this flag has to avoid.
await page.goto('/');
await expect(page.locator('footer a[href="/about"]')).toHaveCount(0);
// And a sitemap must never advertise a URL that 404s.
const sitemap = await page.request.get('/content-sitemap.xml');
expect(await sitemap.text()).not.toContain('/about');
});
test('with the blog off, it is absent rather than empty', async ({ page }) => {
test.skip(await blogIsOn(page), 'the blog flag is on in this environment');
expect((await page.request.get('/blog')).status()).toBe(404);
expect((await page.request.get('/blog/rss.xml')).status()).toBe(404);
await page.goto('/');
await expect(page.locator('footer a[href="/blog"]')).toHaveCount(0);
const sitemap = await page.request.get('/content-sitemap.xml');
expect(await sitemap.text()).not.toContain('/blog');
// The admin panel is deliberately NOT flagged: posts have to be writable
// before the blog is readable, or there is nothing to turn on.
expect((await page.request.get('/admin')).status()).not.toBe(404);
});
test('the about page names a human author and is reachable from the footer', async ({ page }) => {
test.skip(!(await aboutPageIsOn(page)), 'the about_page flag is off in this environment');
await page.goto('/');
const aboutLink = page.locator('footer a[href="/about"]');
await expect(aboutLink).toBeVisible();
await aboutLink.click();
await page.waitForURL(/\/about$/);
await expect(page.getByRole('heading', { level: 1 })).toContainText('Tudor');
await expect(page.locator('img[alt*="Tudor"]')).toBeVisible();
// The credibility claim is lived experience plus stated provenance, not
// expertise. If this sentence ever disappears the positioning has drifted.
await expect(page.getByText(/not an education expert/i)).toBeVisible();
const jsonLd = await page
.locator('script[type="application/ld+json"]')
.first()
.textContent();
expect(jsonLd).toContain('"Person"');
// First name only — a surname here would be the one place it leaks.
expect(jsonLd).not.toMatch(/familyName/);
});
test('the blog lists posts and each one renders with a byline', async ({ page }) => {
test.skip(!(await blogIsOn(page)), 'the blog flag is off in this environment');
await page.goto('/blog');
await expect(page.getByRole('heading', { level: 1 })).toBeVisible();
const postLinks = page.locator('a[href^="/blog/"]');
// Data invariant: staging must carry at least one published post. If this
// fails, the environment has no content rather than the code being broken.
expect(await postLinks.count()).toBeGreaterThan(0);
await postLinks.first().click();
await page.waitForURL(/\/blog\/.+/);
await expect(page.getByRole('heading', { level: 1 })).toBeVisible();
await expect(page.getByText(/^By Tudor/)).toBeVisible();
const jsonLd = await page
.locator('script[type="application/ld+json"]')
.first()
.textContent();
expect(jsonLd).toContain('"BlogPosting"');
});
test('the admin panel is not indexable', async ({ page }) => {
const response = await page.request.get('/admin');
expect(response.headers()['x-robots-tag']).toContain('noindex');
});
test('the content sitemap lists the about page and is advertised in robots', async ({ page }) => {
const sitemap = await page.request.get('/content-sitemap.xml');
// Served whatever the flags say: robots.txt names it unconditionally, and
// with both dark it is a valid empty urlset rather than a 404.
expect(sitemap.ok()).toBeTruthy();
if (await aboutPageIsOn(page)) {
expect(await sitemap.text()).toContain('/about');
}
// The school corpus sitemap is proxied from FastAPI; this one is Next's.
// robots.txt must advertise both or the blog never gets discovered.
const robots = await page.request.get('/robots.txt');
const body = await robots.text();
expect(body).toContain('/sitemap.xml');
expect(body).toContain('/content-sitemap.xml');
});
+12 -4
View File
@@ -1,8 +1,16 @@
# API Configuration
NEXT_PUBLIC_API_URL=http://localhost:8000/api
# Browser requests use the same-origin Next.js proxy.
NEXT_PUBLIC_API_URL=/api
# Production API URL (for deployment)
# NEXT_PUBLIC_API_URL=https://api.schoolcompare.co.uk/api
# Absolute URL for server-side fetching and the proxy; include /api.
# In the managed container network this is http://backend:80/api (staging differs).
FASTAPI_URL=http://localhost:8000/api
# Payload CMS runtime configuration. Use the managed environment's database;
# Payload owns the payload schema, independently of the school marts.
DATABASE_URL=postgresql://schoolcompare:CHANGE_THIS_PASSWORD@localhost:5432/schoolcompare
# Generate a secret: python -c "import secrets; print(secrets.token_urlsafe(32))"
# Use distinct secrets for staging and production.
PAYLOAD_SECRET=CHANGE_THIS_TO_A_SECURE_RANDOM_SECRET
# Node Environment
NODE_ENV=development
+1
View File
@@ -39,3 +39,4 @@ yarn-error.log*
# typescript
*.tsbuildinfo
next-env.d.ts
+12 -288
View File
@@ -1,291 +1,15 @@
# Deployment Guide
# Frontend deployment
This guide covers deployment options for the SchoolCompare Next.js application.
Next.js and Payload run in the same frontend container. The maintained deployment
procedure is [docs/DEPLOY.md](../docs/DEPLOY.md), with the production and staging
Portainer compose files at the repository root.
## Deployment Options
The frontend Dockerfile builds a standalone Next.js image. Runtime configuration
supplies `FASTAPI_URL`, `DATABASE_URL` and `PAYLOAD_SECRET`; uploaded CMS media is
persisted in a volume. Promote the built image through the repository's Gitea
workflow after human staging approval.
### Option 1: Vercel (Recommended for Next.js)
Vercel is the easiest and most optimized platform for Next.js applications.
#### Steps:
1. **Install Vercel CLI**:
```bash
npm install -g vercel
```
2. **Login to Vercel**:
```bash
vercel login
```
3. **Deploy**:
```bash
vercel --prod
```
4. **Configure Environment Variables** in Vercel dashboard:
- `NEXT_PUBLIC_API_URL`: Your FastAPI endpoint (e.g., `https://api.schoolcompare.co.uk/api`)
- `FASTAPI_URL`: Same as above for server-side requests
#### Benefits:
- Automatic HTTPS
- Global CDN
- Zero-config deployment
- Automatic preview deployments
- Built-in analytics
---
### Option 2: Docker (Self-hosted)
Deploy using Docker containers for full control.
#### Prerequisites:
- Docker 20+
- Docker Compose 2+
#### Steps:
1. **Build Docker Image**:
```bash
docker build -t schoolcompare-nextjs:latest .
```
2. **Run with Docker Compose**:
```bash
# Create .env file with production variables
echo "NEXT_PUBLIC_API_URL=https://api.schoolcompare.co.uk/api" > .env
echo "FASTAPI_URL=http://backend:8000/api" >> .env
# Start services
docker-compose up -d
```
3. **Verify Deployment**:
```bash
curl http://localhost:3000
```
#### Environment Variables:
- `NEXT_PUBLIC_API_URL`: Public API endpoint (client-side)
- `FASTAPI_URL`: Internal API endpoint (server-side)
- `NODE_ENV`: `production`
---
### Option 3: PM2 (Node.js Process Manager)
Deploy directly on a Node.js server using PM2.
#### Prerequisites:
- Node.js 24+
- PM2 (`npm install -g pm2`)
#### Steps:
1. **Build Application**:
```bash
npm run build
```
2. **Create PM2 Ecosystem File** (`ecosystem.config.js`):
```javascript
module.exports = {
apps: [{
name: 'schoolcompare-nextjs',
script: 'npm',
args: 'start',
cwd: '/path/to/nextjs-app',
instances: 'max',
exec_mode: 'cluster',
env: {
NODE_ENV: 'production',
PORT: 3000,
NEXT_PUBLIC_API_URL: 'https://api.schoolcompare.co.uk/api',
FASTAPI_URL: 'http://localhost:8000/api',
},
}],
};
```
3. **Start with PM2**:
```bash
pm2 start ecosystem.config.js
pm2 save
pm2 startup
```
---
### Option 4: Nginx Reverse Proxy
Use Nginx as a reverse proxy in front of Next.js.
#### Nginx Configuration:
```nginx
server {
listen 80;
server_name schoolcompare.co.uk;
# Redirect to HTTPS
return 301 https://$server_name$request_uri;
}
server {
listen 443 ssl http2;
server_name schoolcompare.co.uk;
# SSL Configuration
ssl_certificate /etc/ssl/certs/schoolcompare.crt;
ssl_certificate_key /etc/ssl/private/schoolcompare.key;
# Security Headers
# frame-ancestors replaces X-Frame-Options so the analytics subdomain
# (Umami heatmap/recorder) can embed the site in an iframe.
add_header Content-Security-Policy "frame-ancestors 'self' https://analytics.schoolcompare.co.uk" always;
add_header X-Content-Type-Options "nosniff" always;
add_header X-XSS-Protection "1; mode=block" always;
# Proxy to Next.js
location / {
proxy_pass http://localhost:3000;
proxy_http_version 1.1;
proxy_set_header Upgrade $http_upgrade;
proxy_set_header Connection 'upgrade';
proxy_set_header Host $host;
proxy_cache_bypass $http_upgrade;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto $scheme;
}
# Proxy to FastAPI
location /api/ {
proxy_pass http://localhost:8000;
proxy_http_version 1.1;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto $scheme;
}
# Cache static files
location /_next/static/ {
proxy_pass http://localhost:3000;
add_header Cache-Control "public, max-age=31536000, immutable";
}
}
```
---
## Pre-Deployment Checklist
- [ ] Run `npm run build` successfully
- [ ] Run `npm test` - all tests pass
- [ ] Environment variables configured
- [ ] FastAPI backend accessible
- [ ] Database migrations applied
- [ ] SSL certificates configured (production)
- [ ] Domain DNS configured
- [ ] Monitoring/logging set up
- [ ] Backup strategy in place
---
## Post-Deployment Verification
1. **Health Check**:
```bash
curl https://schoolcompare.co.uk
```
2. **Test Routes**:
- Home: `https://schoolcompare.co.uk/`
- School Page: `https://schoolcompare.co.uk/school/100001`
- Compare: `https://schoolcompare.co.uk/compare`
- Rankings: `https://schoolcompare.co.uk/rankings`
3. **Check SEO**:
- Sitemap: `https://schoolcompare.co.uk/sitemap.xml`
- Robots: `https://schoolcompare.co.uk/robots.txt`
4. **Performance Audit**:
- Run Lighthouse in Chrome DevTools
- Target scores: 90+ for Performance, Accessibility, Best Practices, SEO
---
## Monitoring
### Recommended Tools:
- **Vercel Analytics** (if using Vercel)
- **Sentry** for error tracking
- **Google Analytics** for user analytics
- **Uptime Robot** for uptime monitoring
### Health Check Endpoint:
The application automatically serves health data at the root route.
---
## Rollback Procedure
### Vercel:
```bash
vercel rollback
```
### Docker:
```bash
docker-compose down
docker-compose up -d --force-recreate
```
### PM2:
```bash
pm2 stop schoolcompare-nextjs
# Restore previous build
pm2 start schoolcompare-nextjs
```
---
## Troubleshooting
### Issue: API requests failing
- **Solution**: Check `NEXT_PUBLIC_API_URL` and `FASTAPI_URL` environment variables
- **Verify**: FastAPI backend is accessible from Next.js container/server
### Issue: Build fails
- **Solution**: Check Node.js version (requires 24+)
- **Clear cache**: `rm -rf .next node_modules && npm install && npm run build`
### Issue: Slow page loads
- **Solution**: Enable caching in API calls
- **Check**: Network latency to FastAPI backend
- **Verify**: CDN is serving static assets
---
## Security Considerations
- ✅ HTTPS enabled
- ✅ Security headers configured (X-Frame-Options, CSP, etc.)
- ✅ API keys in environment variables (never in code)
- ✅ CORS properly configured
- ✅ Rate limiting on API endpoints
- ✅ Regular security updates
- ✅ Dependency vulnerability scanning
---
## Support
For deployment issues, contact the DevOps team or refer to:
- [Next.js Deployment Docs](https://nextjs.org/docs/deployment)
- [Vercel Documentation](https://vercel.com/docs)
- [Docker Documentation](https://docs.docker.com/)
Earlier Vercel and standalone deployment recipes have been retired from this file
because they do not describe the current CMS, persistence and promotion setup.
See [development](../docs/DEVELOPMENT.md) for checks and
[publishing](docs/PUBLISHING.md) for CMS operations.
+7
View File
@@ -53,6 +53,13 @@ COPY --from=builder /app/.next/static ./.next/static
# a miss here is a silent 500 on /opengraph-image, not a build failure.
COPY --from=builder /app/assets ./assets
# Payload writes uploads here, and the compose file mounts a named volume over
# it. The directory must exist and be owned by the runtime user BEFORE the
# mount: Docker seeds a fresh named volume from the image path, so a missing or
# root-owned directory here makes every upload fail with EACCES at runtime,
# long after the build passed. The chown below covers it.
RUN mkdir -p /app/media
# Set correct permissions
RUN chown -R nextjs:nodejs /app
+43 -141
View File
@@ -1,156 +1,58 @@
# SchoolCompare Next.js Application
# SchoolCompare frontend and CMS
Modern Next.js application for comparing primary school KS2 performance across England.
Next.js App Router with React, TypeScript, CSS Modules, Chart.js, Leaflet and
Payload CMS. It serves school search, comparisons, rankings, school/place detail
pages and editorial content across England.
## Features
Start with the [repository overview](../README.md),
[architecture](../docs/ARCHITECTURE.md) and [development checks](../docs/DEVELOPMENT.md).
- **Server-Side Rendering (SSR)**: Fast initial page loads with pre-rendered content
- **Individual School Pages**: Dedicated pages for each school with full SEO optimization
- **Side-by-Side Comparison**: Compare up to 5 schools simultaneously
- **School Rankings**: Top-performing schools by various metrics
- **Interactive Maps**: Leaflet integration for geographic visualization
- **Performance Charts**: Chart.js visualizations for historical data
- **Responsive Design**: Mobile-first approach with full responsive support
- **SEO Optimized**: Dynamic sitemaps, meta tags, and structured data
## Source map
## Tech Stack
| Path | Purpose |
|---|---|
| `app/(frontend)/` | Public root layout, server pages and FastAPI proxy |
| `app/(payload)/` | Payload root layout, `/admin` and `/cms-api` |
| `app/robots.ts`, `app/opengraph-image.tsx`, root icons | Site-wide metadata endpoints |
| `components/` | Client views and reusable display components |
| `components/school/` | School detail sections |
| `lib/api.ts`, `lib/types.ts` | Fetch wrappers and manual school API types |
| `lib/schoolSections.ts`, `lib/compareLogic.ts` | Presentation decisions and data preparation |
| `context/`, `hooks/` | Comparison state, suggestion state and responsive behaviour |
| `collections/`, `blocks/`, `migrations/` | CMS schema and production migrations |
| `__tests__/` | Jest and React Testing Library tests |
- **Framework**: Next.js 16 (App Router)
- **Language**: TypeScript 5
- **Styling**: CSS Modules + CSS Variables
- **State Management**: React Context API + URL state
- **Data Fetching**: SWR (client-side) + Next.js fetch (server-side)
- **Charts**: Chart.js + react-chartjs-2
- **Maps**: Leaflet + react-leaflet
- **Testing**: Jest + React Testing Library
- **Validation**: Zod
Do not introduce a shared `app/layout.tsx`: public pages and Payload have separate
root layouts. Keep root metadata files outside the route groups.
## Getting Started
## Data and state
### Prerequisites
Server pages fetch initial data directly from `FASTAPI_URL`. Browser fetches use
`/api` by default, forwarded by `app/(frontend)/api/[...path]/route.ts`.
`FASTAPI_URL` must include `/api`. See `.env.example` for CMS and API settings.
- Node.js 24+ (using nvm recommended)
- FastAPI backend running on port 8000
State uses React hooks/context, URL search parameters and localStorage for the
comparison basket. SWR is not installed. Maps use dynamic Leaflet wrappers.
Revalidation intervals are configured in fetch wrappers and pages; they vary by
resource. Backend reloads do not automatically invalidate every Next.js cache.
### Installation
## Commands
```bash
# Install dependencies
npm install
# Copy environment variables
cp .env.example .env.local
# Update .env.local with your configuration
```
### Development
```bash
# Start development server
npm run dev
# Open http://localhost:3000
```
### Building
```bash
# Build for production
```sh
npm ci
npm run typecheck
npm test -- --runInBand
npm run build
# Start production server
npm start
```
### Testing
`test:watch` and `test:coverage` are also available. There is no `lint` script.
A running application needs the backend/data environment described in the
[development guide](../docs/DEVELOPMENT.md).
```bash
# Run tests
npm test
After CMS field or editor changes, run `npm run generate:importmap`. Keep
`payload-types.ts` generated from the CMS schema rather than editing it by hand.
The build must work without a database connection; avoid module-scope CMS queries
and DB-backed `generateStaticParams` functions.
# Run tests in watch mode
npm run test:watch
# Run tests with coverage
npm run test:coverage
```
### Linting
```bash
# Run ESLint
npm run lint
```
## Project Structure
```
nextjs-app/
├── app/ # App Router pages
│ ├── layout.tsx # Root layout
│ ├── page.tsx # Home page
│ ├── compare/ # Compare page
│ ├── rankings/ # Rankings page
│ ├── school/[urn]/ # Individual school pages
│ ├── sitemap.ts # Dynamic sitemap
│ └── robots.ts # Robots.txt
├── components/ # React components
│ ├── SchoolCard.tsx # School card component
│ ├── FilterBar.tsx # Search/filter controls
│ ├── ComparisonView.tsx # Comparison interface
│ ├── RankingsView.tsx # Rankings table
│ └── ...
├── lib/ # Utility libraries
│ ├── api.ts # API client
│ ├── types.ts # TypeScript types
│ └── utils.ts # Helper functions
├── hooks/ # Custom React hooks
├── context/ # React Context providers
├── styles/ # Global styles
├── public/ # Static assets
└── __tests__/ # Test files
```
## Environment Variables
| Variable | Description | Default |
|----------|-------------|---------|
| `NEXT_PUBLIC_API_URL` | Public API endpoint (client-side) | `http://localhost:8000/api` |
| `FASTAPI_URL` | Server-side API endpoint | `http://localhost:8000/api` |
| `NODE_ENV` | Environment mode | `development` |
## Performance Optimizations
- **Server-Side Rendering**: Initial HTML rendered on server
- **Static Generation**: Where possible, pages are pre-generated
- **Image Optimization**: Next.js Image component with AVIF/WebP support
- **Code Splitting**: Automatic route-based code splitting
- **Dynamic Imports**: Heavy components loaded on demand
- **API Caching**: Configurable revalidation for data fetching
- **Bundle Optimization**: Tree shaking and minification
- **Compression**: Gzip compression enabled
## SEO Features
- **Dynamic Meta Tags**: Generated per page with Next.js Metadata API
- **Open Graph**: Social media optimization
- **JSON-LD**: Structured data for search engines
- **Sitemap**: Auto-generated from database
- **Robots.txt**: Search engine crawling rules
- **Canonical URLs**: Duplicate content prevention
## Browser Support
- Chrome (latest)
- Firefox (latest)
- Safari (latest)
- Edge (latest)
## License
Proprietary - SchoolCompare
## Support
For issues and questions, please contact the development team.
See [publishing](docs/PUBLISHING.md) for CMS operations and
[deployment](../docs/DEPLOY.md) for staging and production promotion.
@@ -8,7 +8,7 @@
// environment provides — under jsdom this suite fails on import, not on an
// assertion.
import { NextRequest } from 'next/server';
import { GET } from '@/app/api/[...path]/route';
import { GET } from '@/app/(frontend)/api/[...path]/route';
function request(path: string) {
return new NextRequest(`http://localhost:3000/api/${path}`);
@@ -0,0 +1,34 @@
import { metadata } from '@/app/(frontend)/about/page';
import { personJsonLd, organizationJsonLd } from '@/lib/jsonld';
describe('/about metadata', () => {
it('canonicalises to the bare path', () => {
expect(metadata.alternates?.canonical)
.toBe('https://www.schoolcompare.co.uk/about');
});
});
describe('author structured data', () => {
it('describes a Person with a first name and a photo', () => {
const person = personJsonLd();
expect(person['@type']).toBe('Person');
expect(person.name).toBe('Tudor');
expect(person.image).toBe('https://www.schoolcompare.co.uk/brand/tudor.jpg');
expect(person.url).toBe('https://www.schoolcompare.co.uk/about');
});
it('never publishes a surname or an employer', () => {
// Author identity constraint: first name only. A surname here would be
// the one place it leaks, since JSON-LD is machine-read and archived.
const serialised = JSON.stringify(personJsonLd());
expect(serialised).not.toMatch(/familyName|Sitaru/i);
expect(serialised).not.toMatch(/worksFor|affiliation/i);
});
it('describes the site as an Organization the Person authors for', () => {
const org = organizationJsonLd();
expect(org['@type']).toBe('Organization');
expect(org.name).toBe('schoolcompare');
expect(org.url).toBe('https://www.schoolcompare.co.uk');
});
});
@@ -0,0 +1,66 @@
/**
* The blog index imports getCachedPayload, which pulls in Payload — ESM-only,
* and next/jest will not transform node_modules. Mocking that one module keeps
* the page's metadata testable without loading the CMS; the mock is never
* called, because `metadata` is a static export evaluated at import time.
*/
jest.mock('@/lib/payload', () => ({ getCachedPayload: jest.fn() }));
import { metadata } from '@/app/(frontend)/blog/page';
import { blogPostingJsonLd, breadcrumbJsonLd } from '@/lib/jsonld';
const post = {
title: 'What the data cannot tell you',
slug: 'what-the-data-cannot-tell-you',
excerpt: 'Results describe one year group on a handful of days.',
publishedAt: '2026-09-15T00:00:00.000Z',
};
describe('/blog metadata', () => {
it('canonicalises to the bare path', () => {
expect(metadata.alternates?.canonical)
.toBe('https://www.schoolcompare.co.uk/blog');
});
});
describe('BlogPosting structured data', () => {
it('names the same Person entity the about page declares', () => {
// By @id, not by repeating the person: search engines must resolve every
// post and the about page to one author entity, or the site has several.
const ld = blogPostingJsonLd(post, { namedAuthor: true });
expect(ld['@type']).toBe('BlogPosting');
expect(ld.author['@id']).toBe('https://www.schoolcompare.co.uk/about#tudor');
expect(ld.publisher['@id']).toBe('https://www.schoolcompare.co.uk#organization');
});
it('attributes to the organization when the about page is dark', () => {
/*
* The two flags are independent, so blog-on-about-off is a reachable
* state. The Person entity lives at /about#tudor and that URL 404s while
* the flag is dark, so claiming it would declare an author that resolves
* to nothing — worse for the blog's credibility than having no named
* author at all. Attribute to the publisher instead.
*/
const ld = blogPostingJsonLd(post, { namedAuthor: false });
expect(ld.author['@id']).toBe('https://www.schoolcompare.co.uk#organization');
expect(JSON.stringify(ld)).not.toContain('/about');
});
it('carries a self-referencing canonical url and the publish date', () => {
const ld = blogPostingJsonLd(post, { namedAuthor: true });
expect(ld.url).toBe(
'https://www.schoolcompare.co.uk/blog/what-the-data-cannot-tell-you',
);
expect(ld.datePublished).toBe('2026-09-15T00:00:00.000Z');
});
});
describe('breadcrumbs', () => {
it('places the post under the blog index', () => {
const ld = breadcrumbJsonLd(post);
expect(ld.itemListElement[0].item).toBe('https://www.schoolcompare.co.uk/blog');
expect(ld.itemListElement[1].item).toBe(
'https://www.schoolcompare.co.uk/blog/what-the-data-cannot-tell-you',
);
});
});
+43 -4
View File
@@ -1,7 +1,8 @@
import { metadata as homeMetadata } from '@/app/page';
import { metadata as rankingsMetadata } from '@/app/rankings/page';
import { metadata as admissionsMetadata } from '@/app/admissions/page';
import { generateMetadata as compareMetadata } from '@/app/compare/page';
import { metadata as homeMetadata } from '@/app/(frontend)/page';
import { metadata as rankingsMetadata } from '@/app/(frontend)/rankings/page';
import { metadata as admissionsMetadata } from '@/app/(frontend)/admissions/page';
import { generateMetadata as compareMetadata } from '@/app/(frontend)/compare/page';
import { metadata as rootMetadata } from '@/app/(frontend)/layout';
describe('canonical URLs', () => {
it('the homepage canonicalises to the bare root', () => {
@@ -128,3 +129,41 @@ describe('C1 snippet copy', () => {
}
});
});
/**
* The share card must be declared, not inherited.
*
* `app/opengraph-image.tsx` is a metadata file convention, and it does attach
* to routes in the app root segment — `_not-found` gets an og:image from it.
* It does NOT attach to the site's pages, which live in the `(frontend)`
* route group whose own layout is a root layout. Staging served og:title,
* og:description, og:url, og:site_name and og:type and no og:image at all,
* so every link pasted into a chat rendered bare.
*
* The file stays at the app root, because /robots.txt and /icon.png depend on
* it being there. The site's root layout points at the route it generates.
*/
describe('the share card', () => {
it('declares an opengraph image on the site root layout', () => {
// No og:image means every link pasted into a chat renders bare.
const images = rootMetadata.openGraph?.images;
expect(images).toBeTruthy();
expect(JSON.stringify(images)).toContain('/opengraph-image');
});
it('declares a twitter image too', () => {
// twitter.card is summary_large_image. Claiming a large-image card and
// supplying no image is worse than claiming a summary card.
// Metadata['twitter'] is a union and `card` is not on every member, so
// this reads the serialised shape rather than narrowing the type.
const twitter = JSON.stringify(rootMetadata.twitter);
expect(twitter).toContain('summary_large_image');
expect(twitter).toContain('/opengraph-image');
});
it('resolves the card to an absolute url via metadataBase', () => {
// The e2e journey does `new URL(ogUrl)`, which throws on a relative path.
expect(rootMetadata.metadataBase?.toString())
.toBe('https://www.schoolcompare.co.uk/');
});
});
@@ -0,0 +1,66 @@
/**
* next.config.mjs carries the staging noindex rule. Breaking it turns
* stx.schoolcompare.co.uk into a fully crawlable duplicate of production,
* and nothing else in the suite would notice.
*
* The non-null assertions are deliberate: every key asserted here is optional
* on NextConfig, and a missing one is precisely the regression under test, so
* the assertion below should fail the test rather than the compile.
*/
import nextConfig from '@/next.config.mjs';
async function headerRules() {
return nextConfig.headers!();
}
describe('next.config.mjs', () => {
it('keeps the staging host out of the index', async () => {
const headers = await headerRules();
const stagingRule = headers.find((rule) =>
rule.has?.some(
(cond) => cond.type === 'host' && cond.value === 'stx.schoolcompare.co.uk',
),
);
expect(stagingRule).toBeDefined();
expect(stagingRule!.headers).toContainEqual({
key: 'X-Robots-Tag',
value: 'noindex, nofollow',
});
});
it('still emits standalone output for the Docker runner', () => {
expect(nextConfig.output).toBe('standalone');
});
it('still traces the share-card fonts into the standalone bundle', () => {
expect(nextConfig.outputFileTracingIncludes!['/opengraph-image']).toEqual([
'./assets/**',
]);
});
it('still allows the analytics subdomain to frame the site', async () => {
const headers = await headerRules();
const csp = headers
.flatMap((rule) => rule.headers)
.find((header) => header.key === 'Content-Security-Policy');
expect(csp).toBeDefined();
expect(csp!.value).toContain('https://analytics.schoolcompare.co.uk');
});
});
describe('admin surface', () => {
it('serves noindex on the admin panel and the CMS API', async () => {
// robots.txt disallows these too, but a Disallow only blocks crawling — a
// URL found from an external link can still be indexed without ever being
// fetched. This header is what actually keeps them out.
const headers = await headerRules();
for (const source of ['/admin/:path*', '/cms-api/:path*']) {
const rule = headers.find((entry) => entry.source === source);
expect(rule).toBeDefined();
expect(rule!.headers).toContainEqual({
key: 'X-Robots-Tag',
value: 'noindex, nofollow',
});
}
});
});
@@ -1,4 +1,4 @@
import { generateMetadata as placeMeta } from '@/app/schools/[place]/page';
import { generateMetadata as placeMeta } from '@/app/(frontend)/schools/[place]/page';
jest.mock('@/lib/places', () => ({
...jest.requireActual('@/lib/places'),
+20
View File
@@ -0,0 +1,20 @@
import robots from '@/app/robots';
describe('robots.txt', () => {
it('disallows the admin panel and the CMS API', () => {
const rules = robots().rules;
const rule = Array.isArray(rules) ? rules[0] : rules;
expect(rule.disallow).toEqual(
expect.arrayContaining(['/api/', '/_next/', '/admin/', '/cms-api/']),
);
});
});
describe('sitemap discovery', () => {
it('lists both the proxied school sitemap and the Next-owned content sitemap', () => {
expect(robots().sitemap).toEqual([
'https://www.schoolcompare.co.uk/sitemap.xml',
'https://www.schoolcompare.co.uk/content-sitemap.xml',
]);
});
});
@@ -0,0 +1,47 @@
/**
* The footer is the only navigational route to /about and /blog, so it is
* where a dark flag would otherwise leave a link into a 404.
*
* Both props default to false. A caller that forgets to pass them hides the
* links, which is the direction that cannot break a page — the same reasoning
* as backend/flags.py's "every flag defaults to False".
*/
import { render, screen } from '@testing-library/react';
import { Footer } from '@/components/Footer';
describe('footer feature links', () => {
it('links to both when both flags are on', () => {
render(<Footer aboutEnabled blogEnabled />);
expect(screen.getByRole('link', { name: /who's behind this/i }))
.toHaveAttribute('href', '/about');
expect(screen.getByRole('link', { name: /^blog$/i }))
.toHaveAttribute('href', '/blog');
});
it('omits the about link when that flag is dark', () => {
render(<Footer blogEnabled />);
expect(screen.queryByRole('link', { name: /who's behind this/i })).toBeNull();
expect(screen.getByRole('link', { name: /^blog$/i })).toBeInTheDocument();
});
it('omits the blog link when that flag is dark', () => {
render(<Footer aboutEnabled />);
expect(screen.queryByRole('link', { name: /^blog$/i })).toBeNull();
expect(screen.getByRole('link', { name: /who's behind this/i }))
.toBeInTheDocument();
});
it('drops the whole section when both are dark, not an empty heading', () => {
// Shipping dark means the footer renders as it did before the feature
// existed, not as a section with its contents removed.
render(<Footer />);
expect(screen.queryByRole('heading', { name: /^about$/i })).toBeNull();
expect(screen.queryByRole('link', { name: /who's behind this/i })).toBeNull();
expect(screen.queryByRole('link', { name: /^blog$/i })).toBeNull();
});
it('defaults to dark when a caller passes nothing', () => {
render(<Footer />);
expect(screen.queryByRole('link', { name: /who's behind this/i })).toBeNull();
});
});
@@ -0,0 +1,92 @@
/**
* The module that ends the stranding: before it, a school page's only anchor
* pointed at the school's own website, so ~27k pages sent authority off-site
* and none of it reached the location layer.
*/
import { render, screen } from '@testing-library/react';
import { NearbyPlaces } from '@/components/school/NearbyPlaces';
const essex = { kind: 'authority', slug: 'essex', name: 'Essex', count: 480, url: '/schools/authority/essex', phases: [] };
const brentwood = { kind: 'town', slug: 'brentwood', name: 'Brentwood', count: 37, url: '/schools/brentwood', phases: [] };
const cm15 = { kind: 'outcode', slug: 'cm15', name: 'CM15', count: 12, url: '/schools/near/cm15', phases: [] };
describe('NearbyPlaces', () => {
it('links to every place the school belongs to', () => {
render(<NearbyPlaces places={[essex, brentwood, cm15]} />);
expect(screen.getByRole('link', { name: /Brentwood/ }))
.toHaveAttribute('href', '/schools/brentwood');
expect(screen.getByRole('link', { name: /Essex/ }))
.toHaveAttribute('href', '/schools/authority/essex');
expect(screen.getByRole('link', { name: /CM15/ }))
.toHaveAttribute('href', '/schools/near/cm15');
});
it('says how many schools each link leads to', () => {
// An anchor that states its destination's size is worth more to a reader
// and to a crawler than "see more".
render(<NearbyPlaces places={[brentwood]} />);
expect(screen.getByRole('link', { name: /37 schools in Brentwood/ }))
.toBeInTheDocument();
});
it('renders nothing at all when the school has no published places', () => {
// Not an empty heading. A school whose town and authority both fall below
// the threshold has nowhere to point, and the page should look as it did
// before the module existed.
const { container } = render(<NearbyPlaces places={[]} />);
expect(container).toBeEmptyDOMElement();
});
it('puts the narrowest place first, which is the most useful link', () => {
// The API orders widest-first for the breadcrumb; a reader on a school
// page wants its town before its county.
render(<NearbyPlaces places={[essex, brentwood, cm15]} />);
const hrefs = screen.getAllByRole('link').map((a) => a.getAttribute('href'));
expect(hrefs.indexOf('/schools/brentwood'))
.toBeLessThan(hrefs.indexOf('/schools/authority/essex'));
});
it('handles a singular count without saying "1 schools"', () => {
render(<NearbyPlaces places={[{ ...brentwood, count: 1 }]} />);
expect(screen.getByRole('link', { name: /1 school in Brentwood/ }))
.toBeInTheDocument();
});
it('links the phase page the school appears on', () => {
// "primary schools in brentwood" is the query these pages exist for.
render(<NearbyPlaces places={[{
...brentwood,
phases: [{ phase: 'primary', count: 22, url: '/schools/brentwood/primary' }],
}]} />);
expect(screen.getByRole('link', { name: /22 primary schools in Brentwood/ }))
.toHaveAttribute('href', '/schools/brentwood/primary');
});
it('links both phase pages for an all-through school', () => {
render(<NearbyPlaces places={[{
...brentwood,
phases: [
{ phase: 'primary', count: 22, url: '/schools/brentwood/primary' },
{ phase: 'secondary', count: 9, url: '/schools/brentwood/secondary' },
],
}]} />);
expect(screen.getByRole('link', { name: /22 primary schools/ })).toBeInTheDocument();
expect(screen.getByRole('link', { name: /9 secondary schools/ })).toBeInTheDocument();
});
it('keeps a phase link next to the place it belongs to', () => {
// Grouping matters: "22 primary schools in Brentwood" directly after
// "37 schools in Brentwood" reads as one place, not two unrelated links.
render(<NearbyPlaces places={[essex, {
...brentwood,
phases: [{ phase: 'primary', count: 22, url: '/schools/brentwood/primary' }],
}]} />);
const hrefs = screen.getAllByRole('link').map((a) => a.getAttribute('href'));
expect(hrefs.indexOf('/schools/brentwood/primary'))
.toBe(hrefs.indexOf('/schools/brentwood') + 1);
});
});
@@ -109,7 +109,7 @@ describe('dark-theme safety', () => {
* simply missed.
*/
describe('third-party surfaces under themed text', () => {
const GLOBALS = path.join(__dirname, '..', '..', 'app', 'globals.css');
const GLOBALS = path.join(__dirname, '..', '..', 'app', '(frontend)', 'globals.css');
/** Leaflet surfaces our own code writes token-coloured text onto. */
const LEAFLET_POPUP_SURFACES = [
@@ -171,7 +171,7 @@ describe('third-party surfaces under themed text', () => {
*/
describe('destination tokens', () => {
const css = fs.readFileSync(
path.join(__dirname, '..', '..', 'app', 'globals.css'), 'utf8');
path.join(__dirname, '..', '..', 'app', '(frontend)', 'globals.css'), 'utf8');
const TOKENS = [
'--dest-sixthform', '--dest-sfcollege', '--dest-fecollege',
+25 -1
View File
@@ -1,4 +1,4 @@
import { getFlags } from '@/lib/flags';
import { getFlags, FLAGS_REVALIDATE } from '@/lib/flags';
// jsdom provides no global fetch, so there is nothing for jest.spyOn to attach
// to — assign it and restore the original afterwards. This is the first test
@@ -31,4 +31,28 @@ describe('getFlags', () => {
mockFetch(async () => ({ ok: false, status: 503 }));
await expect(getFlags()).resolves.toEqual({});
});
/*
* Reading a flag pins the calling route's ISR floor: Next uses the LOWEST
* revalidate among a route's fetches for the whole route. That is why the
* revalidate is an argument rather than the constant.
*
* Every SEO route here declares `revalidate = 604800`. A gate that read
* flags at the 300s default would drop the whole school and place corpus
* from a weekly cache to a 5-minute one, which is a large origin-load
* regression to pay for a feature flag.
*/
it('reads at the 300s floor by default', async () => {
mockFetch(async () => ({ ok: true, json: async () => ({}) }));
await getFlags();
expect((global.fetch as jest.Mock).mock.calls[0][1])
.toEqual({ next: { revalidate: FLAGS_REVALIDATE } });
});
it('lets a caller pass its own route floor instead', async () => {
mockFetch(async () => ({ ok: true, json: async () => ({}) }));
await getFlags(604800);
expect((global.fetch as jest.Mock).mock.calls[0][1])
.toEqual({ next: { revalidate: 604800 } });
});
});
@@ -0,0 +1,68 @@
/**
* School pages had no BreadcrumbList and no links into the location layer.
* Both are fixed by the same data — the `places` array the API now returns —
* so they are tested together.
*/
import { schoolBreadcrumbJsonLd } from '@/lib/jsonld';
const essex = { kind: 'authority', slug: 'essex', name: 'Essex', count: 480, url: '/schools/authority/essex', phases: [] };
const brentwood = { kind: 'town', slug: 'brentwood', name: 'Brentwood', count: 37, url: '/schools/brentwood', phases: [] };
const outcode = { kind: 'outcode', slug: 'cm15', name: 'CM15', count: 12, url: '/schools/near/cm15', phases: [] };
describe('school breadcrumbs', () => {
it('reads home to authority to town to school', () => {
const ld = schoolBreadcrumbJsonLd({
name: 'Brentwood School', url: '/school/100000-brentwood-school',
places: [essex, brentwood],
});
expect(ld['@type']).toBe('BreadcrumbList');
expect(ld.itemListElement.map((i) => i.name))
.toEqual(['schoolcompare', 'Essex', 'Brentwood', 'Brentwood School']);
expect(ld.itemListElement.map((i) => i.position)).toEqual([1, 2, 3, 4]);
});
it('skips a level the school has no published place for', () => {
// A school whose town falls below the publish threshold has no town page.
// The trail closes over the gap rather than linking to a 404.
const ld = schoolBreadcrumbJsonLd({
name: 'Lone School', url: '/school/1-lone-school', places: [essex],
});
expect(ld.itemListElement.map((i) => i.name))
.toEqual(['schoolcompare', 'Essex', 'Lone School']);
expect(ld.itemListElement.map((i) => i.position)).toEqual([1, 2, 3]);
});
it('omits outcodes, which are not a place a breadcrumb reads through', () => {
// CM15 is a useful link in the module but nonsense in a trail: nobody
// navigates Essex → CM15 → school.
const ld = schoolBreadcrumbJsonLd({
name: 'Brentwood School', url: '/school/100000-brentwood-school',
places: [essex, brentwood, outcode],
});
expect(JSON.stringify(ld)).not.toContain('cm15');
});
it('still produces a valid trail when the school has no places at all', () => {
const ld = schoolBreadcrumbJsonLd({
name: 'Orphan School', url: '/school/2-orphan-school', places: [],
});
expect(ld.itemListElement.map((i) => i.name)).toEqual(['schoolcompare', 'Orphan School']);
});
it('uses absolute urls, as every other entity on the site does', () => {
const ld = schoolBreadcrumbJsonLd({
name: 'Brentwood School', url: '/school/100000-brentwood-school',
places: [essex, brentwood],
});
for (const item of ld.itemListElement) {
expect(item.item).toMatch(/^https:\/\/www\.schoolcompare\.co\.uk\//);
}
// The root is the homepage: there is no /schools index page to link to.
expect(ld.itemListElement[0].item).toBe('https://www.schoolcompare.co.uk/');
});
});
@@ -0,0 +1,70 @@
/**
* Payload is ESM-only and next/jest will not transform it, so the collections
* cannot be imported and their sanitised config inspected here (see
* lib/payloadRoutes.ts for the full reasoning). These assert the source of the
* collection definitions instead — enough to catch the settings whose loss is
* silent, and cheap. Behaviour is proved by the e2e journeys against staging.
*/
import fs from 'fs';
import path from 'path';
const read = (file: string) =>
fs.readFileSync(path.join(__dirname, '..', '..', 'collections', file), 'utf8');
const POSTS = read('Posts.ts');
const MEDIA = read('Media.ts');
const CONFIG = fs.readFileSync(
path.join(__dirname, '..', '..', 'payload.config.ts'),
'utf8',
);
describe('posts collection', () => {
it('supports drafts, so saving is not publishing', () => {
expect(POSTS).toMatch(/drafts:\s*true/);
});
it('has a unique, indexed slug for stable URLs', () => {
const slugField = POSTS.slice(POSTS.indexOf("name: 'slug'"));
expect(slugField).toMatch(/unique:\s*true/);
expect(slugField).toMatch(/index:\s*true/);
});
it('hides drafts from anonymous readers at the access layer', () => {
// Payload's docs are explicit: "The `draft` argument alone does not
// restrict documents with _status: 'draft' from being returned by the
// API." The blog pages' where-clause is not enforcement — a direct GET
// /cms-api/posts would return unpublished drafts to anyone. Access
// control returning a query constraint is the only thing that stops it.
expect(POSTS).toMatch(/_status:\s*\{\s*equals:\s*'published'\s*\}/);
expect(POSTS).toMatch(/if\s*\(req\.user\)\s*return true/);
});
it('revalidates the post page when a post changes or is deleted', () => {
// /blog/[slug] is ISR — generated on first request and cached — so an edit
// to an already-published post would otherwise not appear until the
// revalidate window expired, up to an hour of a writer concluding that
// saving is broken. The index and feeds are force-dynamic and need no hook.
expect(POSTS).toContain('afterChange');
expect(POSTS).toContain('afterDelete');
expect(POSTS).toMatch(/revalidatePath\(`\/blog\/\$\{[^}]+\}`\)/);
});
});
describe('media collection', () => {
it('writes uploads to the mounted volume, by absolute path', () => {
// Must match the payload_media mount in docker-compose.portainer.yml.
// Payload 3 requires staticDir to be absolute.
expect(MEDIA).toMatch(/staticDir:\s*'\/app\/media'/);
});
it('requires alt text on every upload', () => {
const altField = MEDIA.slice(MEDIA.indexOf("name: 'alt'"));
expect(altField).toMatch(/required:\s*true/);
});
});
describe('payload config', () => {
it('registers every collection', () => {
expect(CONFIG).toMatch(/collections:\s*\[Users,\s*Posts,\s*Media\]/);
});
});
@@ -0,0 +1,67 @@
/**
* The admin panel does not import field components directly. Payload sends the
* client a *path* for each one — a richText field's is
* `@payloadcms/richtext-lexical/rsc#RscEntryLexicalField` — and resolves it
* through this generated map. An entry that is missing from the map is not an
* error the panel reports: the field simply does not render.
*
* That failure is quietly awful, because `required: true` is enforced on the
* server regardless. A writer gets a new-post form with no Content editor and
* a save that refuses on a field they were never shown.
*
* The map is generated by `npx payload generate:importmap`, so it drifts every
* time a field or a lexical feature is added and nobody re-runs it. These
* assert the entries the current config needs.
*/
import fs from 'fs';
import path from 'path';
const MAP = fs.readFileSync(
path.join(__dirname, '..', '..', 'app', '(payload)', 'admin', 'importMap.js'),
'utf8',
);
const POSTS = fs.readFileSync(
path.join(__dirname, '..', '..', 'collections', 'Posts.ts'),
'utf8',
);
describe('admin import map', () => {
it('resolves the richText field, so Content renders in the editor', () => {
// Guarded because Posts.content is required: without this entry the field
// is invisible and the post is unsaveable.
expect(POSTS).toMatch(/type:\s*'richText'/);
expect(MAP).toContain('@payloadcms/richtext-lexical/rsc#RscEntryLexicalField');
});
it('resolves the richText cell, so the list view can render the column', () => {
expect(MAP).toContain('@payloadcms/richtext-lexical/rsc#RscEntryLexicalCell');
});
it('resolves the diff component, which the drafts UI needs', () => {
// versions.drafts is on, so the panel offers version comparison.
expect(POSTS).toMatch(/drafts:\s*true/);
expect(MAP).toContain('@payloadcms/richtext-lexical/rsc#LexicalDiffComponent');
});
it('resolves BlocksFeature, so the Callout block is insertable', () => {
expect(POSTS).toContain('BlocksFeature');
expect(MAP).toContain('@payloadcms/richtext-lexical/client#BlocksFeatureClient');
});
it('resolves the default toolbar features the editor is built with', () => {
// defaultFeatures is spread into the editor config; each one contributes a
// client component the toolbar cannot render without.
for (const feature of [
'BoldFeatureClient',
'ItalicFeatureClient',
'HeadingFeatureClient',
'LinkFeatureClient',
'UploadFeatureClient',
'UnorderedListFeatureClient',
'OrderedListFeatureClient',
'InlineToolbarFeatureClient',
]) {
expect(MAP).toContain(`@payloadcms/richtext-lexical/client#${feature}`);
}
});
});
@@ -0,0 +1,52 @@
/**
* The generated migration is schema-qualified to "payload" throughout but does
* not create that schema — `schemaName` says where tables go, it does not
* create anything. On staging and production, which have never run it, the
* whole migration fails with `schema "payload" does not exist`.
*
* The CREATE SCHEMA is therefore hand-added, which makes it exactly the kind
* of edit a regeneration silently discards. This is the guard.
*/
import fs from 'fs';
import path from 'path';
const DIR = path.join(__dirname, '..', '..', 'migrations');
function migrationFiles() {
return fs
.readdirSync(DIR)
.filter((f) => f.endsWith('.ts') && f !== 'index.ts');
}
describe('payload migrations', () => {
it('ships at least one migration, so a container has tables to find', () => {
expect(migrationFiles().length).toBeGreaterThan(0);
});
it('creates the payload schema before creating anything in it', () => {
const initial = migrationFiles().find((f) => f.includes('initial'))!;
const sql = fs.readFileSync(path.join(DIR, initial), 'utf8');
expect(sql).toMatch(/CREATE SCHEMA IF NOT EXISTS "payload"/);
// Ordering matters: the schema must be created before the first object
// that lives in it, or the migration fails on its first statement.
expect(sql.indexOf('CREATE SCHEMA IF NOT EXISTS "payload"'))
.toBeLessThan(sql.indexOf('CREATE TABLE "payload"'));
});
it('creates the tables the app queries on boot', () => {
const initial = migrationFiles().find((f) => f.includes('initial'))!;
const sql = fs.readFileSync(path.join(DIR, initial), 'utf8');
for (const table of ['users', 'posts', '_posts_v', 'media', 'payload_migrations']) {
expect(sql).toContain(`CREATE TABLE "payload"."${table}"`);
}
});
it('is wired into the adapter, so it runs on server init', () => {
const config = fs.readFileSync(
path.join(__dirname, '..', '..', 'payload.config.ts'), 'utf8',
);
expect(config).toMatch(/prodMigrations:\s*migrations/);
});
});
@@ -0,0 +1,44 @@
/**
* Guards the one thing about Payload's mounting that fails silently.
*
* payload.config.ts itself cannot be imported here — Payload is ESM-only and
* next/jest will not transform it — so this asserts the shared constants and
* that the config actually wires them in, by reading its source. The live
* proof that /api still reaches FastAPI is the e2e journeys, which call
* /api/schools against the running app.
*/
import fs from 'fs';
import path from 'path';
import { PAYLOAD_API_ROUTE, PAYLOAD_ADMIN_ROUTE } from '@/lib/payloadRoutes';
const CONFIG = fs.readFileSync(
path.join(__dirname, '..', '..', 'payload.config.ts'),
'utf8',
);
describe('payload mount points', () => {
it('serves the CMS API from /cms-api, never /api', () => {
// /api is the FastAPI proxy's catch-all. Payload's default would be
// swallowed by it and forwarded to the backend, silently.
expect(PAYLOAD_API_ROUTE).toBe('/cms-api');
expect(PAYLOAD_API_ROUTE).not.toBe('/api');
});
it('serves the admin panel from /admin', () => {
expect(PAYLOAD_ADMIN_ROUTE).toBe('/admin');
});
it('wires both constants into the Payload config', () => {
expect(CONFIG).toContain('PAYLOAD_API_ROUTE');
expect(CONFIG).toContain('PAYLOAD_ADMIN_ROUTE');
});
it('never hardcodes a routes block that could drift from the constants', () => {
expect(CONFIG).not.toMatch(/routes:\s*\{[^}]*api:\s*['"]/);
});
it('isolates CMS tables in their own postgres schema', () => {
// Blog content must stay separate from school marts and Airflow metadata.
expect(CONFIG).toMatch(/schemaName:\s*['"]payload['"]/);
});
});
@@ -21,7 +21,7 @@ import {
import { nationalAveragesFixture } from './schoolFixtures';
// The shell calls useComparison(), which throws outside the provider. In the
// app this wrapper comes from app/layout.tsx.
// app this wrapper comes from app/(frontend)/layout.tsx.
function withProviders(ui: ReactNode) {
return <ComparisonProvider>{ui}</ComparisonProvider>;
}
@@ -0,0 +1,82 @@
.page {
max-width: 42rem;
margin: 0 auto;
padding: 2.5rem 1.25rem 4rem;
}
.header {
display: flex;
align-items: center;
gap: 1.25rem;
margin-bottom: 2rem;
}
.portrait {
border-radius: 50%;
border: 2px solid var(--border);
object-fit: cover;
flex-shrink: 0;
}
.kicker {
font-family: var(--font-ui);
font-size: 0.75rem;
font-weight: 600;
text-transform: uppercase;
letter-spacing: 0.06em;
color: var(--brand);
margin: 0 0 0.35rem;
}
.heading {
font-family: var(--font-display);
font-size: clamp(1.5rem, 4vw, 2rem);
font-weight: 700;
line-height: 1.2;
color: var(--text-primary);
margin: 0;
}
.subheading {
font-family: var(--font-display);
font-size: 1.15rem;
font-weight: 600;
color: var(--text-primary);
margin: 2.25rem 0 0.75rem;
}
.prose p {
font-family: var(--font-ui);
font-size: 1rem;
line-height: 1.7;
color: var(--text-secondary);
margin: 0 0 1.1rem;
}
/* The opening paragraph carries the page. Larger, and in the primary ink
rather than the secondary, so it reads as a voice rather than as body copy.
Must stay in the descendant form: `.prose p` scores (0,1,1) and would beat a
bare `.lede` at (0,1,0), so simplifying this selector silently reverts the
lede to ordinary body copy. */
.prose .lede {
font-size: 1.125rem;
color: var(--text-primary);
}
.link {
color: var(--brand);
font-weight: 600;
}
.link:hover {
color: var(--brand-strong);
}
@media (max-width: 480px) {
.header {
flex-direction: column;
align-items: flex-start;
gap: 1rem;
}
}
+143
View File
@@ -0,0 +1,143 @@
import type { Metadata } from 'next';
import Image from 'next/image';
import { notFound } from 'next/navigation';
import { absoluteUrl } from '@/lib/site';
import { getFlags } from '@/lib/flags';
import { personJsonLd, organizationJsonLd } from '@/lib/jsonld';
import styles from './About.module.css';
export const metadata: Metadata = {
title: 'About',
description:
'Who builds schoolcompare, why it exists, and where its numbers come from.',
alternates: { canonical: absoluteUrl('/about') },
};
/*
* Gated on about_page. The default 300s read is the right floor here: this
* page declares no revalidate of its own, so nothing is lost by it, and a flip
* lands within five minutes.
*
* notFound(), not a redirect: while the flag is dark this URL does not exist,
* and a 404 is what tells a crawler not to keep it.
*/
export default async function AboutPage() {
const flags = await getFlags();
if (flags.about_page !== true) notFound();
const jsonLd = {
'@context': 'https://schema.org',
'@graph': [personJsonLd(), organizationJsonLd()],
};
return (
<div className={styles.page}>
<script
type="application/ld+json"
dangerouslySetInnerHTML={{ __html: JSON.stringify(jsonLd) }}
/>
<header className={styles.header}>
<Image
src="/brand/tudor.jpg"
alt="Tudor, who builds schoolcompare"
width={96}
height={96}
className={styles.portrait}
priority
/>
<div>
<p className={styles.kicker}>Who&apos;s behind this</p>
<h1 className={styles.heading}>I&apos;m Tudor. I built this site.</h1>
</div>
</header>
<div className={styles.prose}>
<p className={styles.lede}>
I&apos;m a parent in south-west London. When we started looking at
primary schools, I found the information I needed was all published,
and almost impossible to hold in one place.
</p>
<p>
SATs results were in one government table. Ofsted judgements were in a
separate service, in a format that had just changed. Admissions
distances were buried in council PDFs, a different one per borough,
each with its own layout. I ended up building a spreadsheet, and then
I got tired of the spreadsheet.
</p>
<p>
So I built this instead. It pulls the official figures into one place
and puts them side by side, which is what I wanted and could not find.
</p>
<h2 className={styles.subheading}>I&apos;m not an education expert</h2>
<p>
I want to be straightforward about that. I&apos;m not a teacher, a
governor, or an education researcher. I have no qualification that
makes my opinion about a school worth more than yours.
</p>
<p>
What I do have is the problem itself. I&apos;m going through primary
admissions right now, and I work with data for a living. That
combination is enough to take published figures and present them
honestly. It is not enough to tell you which school is right for your
child, and this site never tries to.
</p>
<h2 className={styles.subheading}>Where the numbers come from</h2>
<p>
Everything here is official published data: Key Stage 2 and Key Stage
4 results and school characteristics from the Department for
Education, inspection outcomes from Ofsted, and admissions data from
local authorities. Nothing is estimated, modelled or filled in. Where
a figure is missing, the page says so rather than showing a guess.
</p>
<p>
This is an independent site. It is not affiliated with the Department
for Education or with Ofsted, and nobody pays to appear on it or to
rank higher.
</p>
<h2 className={styles.subheading}>What the data can&apos;t tell you</h2>
<p>
A school is not its results. The figures here describe one year group,
on a handful of days, measured in a way that suits national statistics
rather than your child. A small cohort makes percentages swing wildly.
In a class of thirty, one pupil is worth more than three points.
Results say nothing at all about whether a child will be happy
somewhere.
</p>
<p>
I try to build that honesty into the site rather than just say it
here. Special schools and pupil referral units are never compared
against a mainstream national average, because that comparison is
meaningless and makes good schools look like failing ones. Where a
number is unreliable, the aim is for the page to tell you before you
draw a conclusion from it.
</p>
<h2 className={styles.subheading}>If something&apos;s wrong</h2>
<p>
Tell me and I&apos;ll fix it. If a figure looks wrong, or a page gives
a misleading impression of a school, I genuinely want to know.
It&apos;s the fastest way this gets better.
</p>
<p>
<a href="mailto:contact@schoolcompare.co.uk" className={styles.link}>
contact@schoolcompare.co.uk
</a>
</p>
</div>
</div>
);
}
@@ -0,0 +1,76 @@
.page {
max-width: 42rem;
margin: 0 auto;
padding: 2.5rem 1.25rem 4rem;
}
.header { margin-bottom: 2.5rem; }
.kicker {
font-family: var(--font-ui);
font-size: 0.75rem;
font-weight: 600;
text-transform: uppercase;
letter-spacing: 0.06em;
color: var(--brand);
margin: 0 0 0.35rem;
}
.heading {
font-family: var(--font-display);
font-size: clamp(1.5rem, 4vw, 2rem);
font-weight: 700;
line-height: 1.2;
color: var(--text-primary);
margin: 0 0 0.75rem;
}
.standfirst {
font-family: var(--font-ui);
font-size: 1.05rem;
line-height: 1.65;
color: var(--text-secondary);
margin: 0;
}
.list { list-style: none; padding: 0; margin: 0; }
.item {
padding: 1.5rem 0;
border-top: 1px solid var(--border);
}
.date {
font-family: var(--font-ui);
font-size: 0.8rem;
color: var(--text-muted);
/* Inter's tabular numerals keep a column of dates aligned. */
font-variant-numeric: tabular-nums;
}
.itemTitle {
font-family: var(--font-display);
font-size: 1.25rem;
font-weight: 600;
line-height: 1.3;
margin: 0.35rem 0 0.5rem;
}
.itemLink { color: var(--text-primary); text-decoration: none; }
.itemLink:hover { color: var(--brand); }
.excerpt {
font-family: var(--font-ui);
font-size: 0.95rem;
line-height: 1.65;
color: var(--text-secondary);
margin: 0;
}
.empty {
font-family: var(--font-ui);
color: var(--text-muted);
}
.link { color: var(--brand); font-weight: 600; }
.link:hover { color: var(--brand-strong); }
@@ -0,0 +1,87 @@
.page {
max-width: 42rem;
margin: 0 auto;
padding: 2.5rem 1.25rem 4rem;
}
.crumb {
font-family: var(--font-ui);
font-size: 0.85rem;
margin-bottom: 1.25rem;
}
.heading {
font-family: var(--font-display);
font-size: clamp(1.6rem, 5vw, 2.25rem);
font-weight: 700;
line-height: 1.2;
color: var(--text-primary);
margin: 0 0 0.75rem;
}
.byline {
font-family: var(--font-ui);
font-size: 0.9rem;
color: var(--text-muted);
margin: 0 0 2rem;
}
.hero {
width: 100%;
height: auto;
border-radius: 10px;
border: 1px solid var(--border);
margin-bottom: 2rem;
}
/* Rich-text output: the editor emits plain elements, so these are styled by
descendant selector rather than by class. */
.prose p {
font-family: var(--font-ui);
font-size: 1rem;
line-height: 1.7;
color: var(--text-secondary);
margin: 0 0 1.1rem;
}
.prose h2 {
font-family: var(--font-display);
font-size: 1.25rem;
font-weight: 600;
color: var(--text-primary);
margin: 2.25rem 0 0.75rem;
}
.prose h3 {
font-family: var(--font-display);
font-size: 1.05rem;
font-weight: 600;
color: var(--text-primary);
margin: 1.75rem 0 0.6rem;
}
.prose ul,
.prose ol {
font-family: var(--font-ui);
font-size: 1rem;
line-height: 1.7;
color: var(--text-secondary);
padding-left: 1.35rem;
margin: 0 0 1.1rem;
}
.prose li { margin-bottom: 0.4rem; }
.prose a { color: var(--brand); font-weight: 500; }
.prose a:hover { color: var(--brand-strong); }
.prose blockquote {
border-left: 3px solid var(--border-strong);
padding-left: 1rem;
margin: 1.5rem 0;
color: var(--text-muted);
font-style: italic;
}
.link { color: var(--brand); font-weight: 600; }
.link:hover { color: var(--brand-strong); }
@@ -0,0 +1,194 @@
import { cache } from 'react';
import type { Metadata } from 'next';
import Link from 'next/link';
import { notFound } from 'next/navigation';
import { RichText } from '@payloadcms/richtext-lexical/react';
import type { JSXConvertersFunction } from '@payloadcms/richtext-lexical/react';
import { getCachedPayload } from '@/lib/payload';
import type { Post, Media } from '@/payload-types';
import { absoluteUrl } from '@/lib/site';
import { getFlags } from '@/lib/flags';
import {
blogPostingJsonLd,
breadcrumbJsonLd,
personJsonLd,
organizationJsonLd,
} from '@/lib/jsonld';
import { CalloutBlock } from '@/components/blog/CalloutBlock';
import styles from './Post.module.css';
/*
* ISR. Unlike the index, this route has a dynamic param and no
* generateStaticParams, so there is nothing for the build to prerender: each
* post is generated on first request and cached until the collection's
* afterChange hook revalidates it. That hook is what makes an edit to an
* already-published post appear immediately.
*/
export const revalidate = 3600;
/**
* heroImage is `number | Media | null`: an id when the query is shallow, the
* populated document at depth 1. Both pages query at depth 1, but narrowing
* rather than asserting keeps it correct if that ever changes.
*/
function heroOf(post: Post): Media | null {
return typeof post.heroImage === 'object' && post.heroImage !== null
? post.heroImage
: null;
}
/**
* Spreads the default converters and adds the one custom block.
*
* Without the spread, every default node type — paragraphs, headings, links —
* loses its renderer and the post body comes out empty.
*/
const calloutConverters: JSXConvertersFunction = ({ defaultConverters }) => ({
...defaultConverters,
blocks: {
// Annotated because the generic block converter cannot infer a custom
// block's field shape; String() guards the values regardless.
callout: ({ node }: { node: { fields: Record<string, unknown> } }) => (
<CalloutBlock
tone={String(node.fields.tone ?? 'caveat')}
body={String(node.fields.body ?? '')}
/>
),
},
});
/**
* Wrapped in React's cache() because Next calls generateMetadata and the page
* component separately for the same request — without it, every post view runs
* this query against Postgres twice. cache() dedupes within a single request
* only, so it never serves one visitor's request from another's.
*/
const findPost = cache(async (slug: string) => {
const payload = await getCachedPayload();
const { docs } = await payload.find({
collection: 'posts',
where: { slug: { equals: slug }, _status: { equals: 'published' } },
limit: 1,
depth: 1,
});
return docs[0] ?? null;
});
function summarise(post: Post) {
return {
title: post.title,
slug: post.slug,
excerpt: post.excerpt,
publishedAt: post.publishedAt,
};
}
export async function generateMetadata(
{ params }: { params: Promise<{ slug: string }> },
): Promise<Metadata> {
const { slug } = await params;
const post = await findPost(slug);
if (!post) return { title: 'Not found' };
const hero = heroOf(post);
return {
title: post.title,
description: post.excerpt,
alternates: { canonical: absoluteUrl(`/blog/${post.slug}`) },
openGraph: {
type: 'article',
title: post.title,
description: post.excerpt,
url: absoluteUrl(`/blog/${post.slug}`),
publishedTime: post.publishedAt,
// A post with a hero image shares that; one without falls through to the
// generated share card at app/opengraph-image.tsx.
...(hero?.url ? { images: [{ url: hero.url }] } : {}),
},
};
}
export default async function PostPage(
{ params }: { params: Promise<{ slug: string }> },
) {
const { slug } = await params;
/*
* Flags read at this route's own declared floor, so gating costs it nothing.
* Checked before the post is fetched: a dark blog should not query Payload.
*/
const flags = await getFlags(3600);
if (flags.blog !== true) notFound();
const namedAuthor = flags.about_page === true;
const post = await findPost(slug);
if (!post) notFound();
const summary = summarise(post);
const hero = heroOf(post);
const jsonLd = {
'@context': 'https://schema.org',
/*
* The Person entity is anchored at /about#tudor, so it is declared only
* when that page exists. Claiming an author whose URL 404s is a worse
* signal than attributing the post to the publisher.
*/
'@graph': [
blogPostingJsonLd(summary, { namedAuthor }),
breadcrumbJsonLd(summary),
...(namedAuthor ? [personJsonLd()] : []),
organizationJsonLd(),
],
};
return (
<article className={styles.page}>
<script
type="application/ld+json"
dangerouslySetInnerHTML={{ __html: JSON.stringify(jsonLd) }}
/>
<nav className={styles.crumb}>
<Link href="/blog" className={styles.link}>Blog</Link>
</nav>
<h1 className={styles.heading}>{summary.title}</h1>
<p className={styles.byline}>
{/* Unlinked while about_page is dark; the flags are independent. */}
By {namedAuthor
? <Link href="/about" className={styles.link}>Tudor</Link>
: 'Tudor'}
{' · '}
<time dateTime={summary.publishedAt}>
{new Date(summary.publishedAt).toLocaleDateString('en-GB', {
day: 'numeric',
month: 'long',
year: 'numeric',
})}
</time>
</p>
{/*
A plain <img>, not next/image: Payload already generated the sized
derivatives on upload (Media's imageSizes), so routing it through the
optimizer would resize an image that is already the right size.
*/}
{hero?.url && (
<img
className={styles.hero}
src={hero.url}
alt={hero.alt ?? ''}
width={hero.width ?? undefined}
height={hero.height ?? undefined}
/>
)}
<div className={styles.prose}>
<RichText data={post.content} converters={calloutConverters} />
</div>
</article>
);
}
+85
View File
@@ -0,0 +1,85 @@
import type { Metadata } from 'next';
import Link from 'next/link';
import { notFound } from 'next/navigation';
import { getCachedPayload } from '@/lib/payload';
import { absoluteUrl } from '@/lib/site';
import { getFlags } from '@/lib/flags';
import styles from './Blog.module.css';
/*
* Dynamic, not ISR.
*
* This route has no dynamic params, so Next prerenders it at build time — and
* CI builds the image with no database reachable, which fails the build. It is
* a single indexed query against Postgres on the same Docker network, so
* rendering per request is cheap, and it means a newly published post appears
* here immediately rather than waiting on a revalidation.
*/
export const dynamic = 'force-dynamic';
export const metadata: Metadata = {
title: 'Blog',
description:
'Notes on what school performance data shows, and what it does not.',
alternates: { canonical: absoluteUrl('/blog') },
};
function formatDate(value: string) {
return new Date(value).toLocaleDateString('en-GB', {
day: 'numeric',
month: 'long',
year: 'numeric',
});
}
export default async function BlogIndexPage() {
const flags = await getFlags();
if (flags.blog !== true) notFound();
const payload = await getCachedPayload();
const { docs } = await payload.find({
collection: 'posts',
where: { _status: { equals: 'published' } },
sort: '-publishedAt',
limit: 50,
depth: 0,
});
return (
<div className={styles.page}>
<header className={styles.header}>
<p className={styles.kicker}>Blog</p>
<h1 className={styles.heading}>Notes on the numbers</h1>
<p className={styles.standfirst}>
What school performance data shows, what it doesn&apos;t, and how to
read it without being misled. Written by{' '}
{/* Plain text when about_page is dark: the two flags are
independent, so this link would otherwise point at a 404. */}
{flags.about_page === true
? <Link href="/about" className={styles.link}>Tudor</Link>
: 'Tudor'}.
</p>
</header>
{docs.length === 0 ? (
<p className={styles.empty}>No posts yet.</p>
) : (
<ul className={styles.list}>
{docs.map((post) => (
<li key={post.id} className={styles.item}>
<time className={styles.date} dateTime={String(post.publishedAt)}>
{formatDate(String(post.publishedAt))}
</time>
<h2 className={styles.itemTitle}>
<Link href={`/blog/${post.slug}`} className={styles.itemLink}>
{post.title}
</Link>
</h2>
<p className={styles.excerpt}>{post.excerpt}</p>
</li>
))}
</ul>
)}
</div>
);
}
@@ -0,0 +1,58 @@
import { getCachedPayload } from '@/lib/payload';
import { absoluteUrl } from '@/lib/site';
import { getFlags } from '@/lib/flags';
/*
* Dynamic, not ISR.
*
* This route has no dynamic params, so Next prerenders it at build time — and
* CI builds the image with no database reachable, which fails the build. It is
* a single indexed query against Postgres on the same Docker network, so
* rendering per request is cheap, and it means a newly published post appears
* here immediately rather than waiting on a revalidation.
*/
export const dynamic = 'force-dynamic';
function escapeXml(value: string): string {
return value.replace(/[<>&'"]/g, (char) =>
({ '<': '&lt;', '>': '&gt;', '&': '&amp;', "'": '&apos;', '"': '&quot;' }[char]!));
}
export async function GET() {
// A dark blog has no feed. 404 rather than an empty channel: an empty feed
// is a live feed with nothing in it, which a reader would keep polling.
const flags = await getFlags();
if (flags.blog !== true) return new Response('Not found', { status: 404 });
const payload = await getCachedPayload();
const { docs } = await payload.find({
collection: 'posts',
where: { _status: { equals: 'published' } },
sort: '-publishedAt',
limit: 50,
depth: 0,
});
const items = docs.map((post) => `
<item>
<title>${escapeXml(String(post.title))}</title>
<link>${absoluteUrl(`/blog/${post.slug}`)}</link>
<guid isPermaLink="true">${absoluteUrl(`/blog/${post.slug}`)}</guid>
<description>${escapeXml(String(post.excerpt))}</description>
<pubDate>${new Date(String(post.publishedAt)).toUTCString()}</pubDate>
</item>`).join('');
const xml = `<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0">
<channel>
<title>schoolcompare blog</title>
<link>${absoluteUrl('/blog')}</link>
<description>What school performance data shows, and what it does not.</description>
<language>en-GB</language>${items}
</channel>
</rss>`;
return new Response(xml, {
headers: { 'Content-Type': 'application/rss+xml; charset=utf-8' },
});
}
File renamed without changes.
@@ -0,0 +1,68 @@
/*
* A second sitemap for the URLs Next owns.
*
* /sitemap.xml is proxied from FastAPI (app/(frontend)/sitemap.xml), which
* knows nothing about Payload — the backend and frontend ship as separate
* images. Rather than teach it, the Next-owned URLs get their own sitemap and
* robots.txt lists both.
*/
import { getCachedPayload } from '@/lib/payload';
import { absoluteUrl } from '@/lib/site';
import { getFlags } from '@/lib/flags';
/*
* Dynamic, not ISR.
*
* This route has no dynamic params, so Next prerenders it at build time — and
* CI builds the image with no database reachable, which fails the build. It is
* a single indexed query against Postgres on the same Docker network, so
* rendering per request is cheap, and it means a newly published post appears
* here immediately rather than waiting on a revalidation.
*/
export const dynamic = 'force-dynamic';
export async function GET() {
/*
* A dark page must not be advertised. Submitting a URL that 404s is the one
* thing a sitemap is not allowed to do, so each entry is gated on the same
* flag that gates the page itself.
*
* With both flags dark this emits a valid, empty <urlset> rather than a 404:
* robots.txt names this sitemap unconditionally, and an empty sitemap is a
* well-formed statement that there is nothing here yet.
*/
const flags = await getFlags();
const aboutEnabled = flags.about_page === true;
const blogEnabled = flags.blog === true;
// Only query Payload when the blog is actually being advertised.
const docs = blogEnabled
? (await (await getCachedPayload()).find({
collection: 'posts',
where: { _status: { equals: 'published' } },
sort: '-publishedAt',
limit: 500,
depth: 0,
})).docs
: [];
const urls: Array<{ loc: string; lastmod: string | null }> = [
...(aboutEnabled ? [{ loc: absoluteUrl('/about'), lastmod: null }] : []),
...(blogEnabled ? [{ loc: absoluteUrl('/blog'), lastmod: null }] : []),
...docs.map((post) => ({
loc: absoluteUrl(`/blog/${post.slug}`),
lastmod: new Date(String(post.updatedAt ?? post.publishedAt)).toISOString(),
})),
];
const xml = `<?xml version="1.0" encoding="UTF-8"?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
${urls.map(({ loc, lastmod }) =>
` <url><loc>${loc}</loc>${lastmod ? `<lastmod>${lastmod}</lastmod>` : ''}</url>`,
).join('\n')}
</urlset>`;
return new Response(xml, {
headers: { 'Content-Type': 'application/xml; charset=utf-8' },
});
}
File renamed without changes.
@@ -7,6 +7,7 @@ import { ComparisonToast } from '@/components/ComparisonToast';
import { RouteTrail } from '@/components/RouteTrail';
import { ComparisonProvider } from '@/context/ComparisonProvider';
import { SITE_URL } from '@/lib/site';
import { getFlags } from '@/lib/flags';
import './globals.css';
// Manrope carries headings and key messaging — the guideline's "friendly,
@@ -58,14 +59,32 @@ export const metadata: Metadata = {
authors: [{ name: 'schoolcompare' }],
manifest: '/manifest.json',
// No `icons` key on purpose: setting it here would override the file
// conventions. app/icon.svg and app/apple-icon.tsx are the source, and
// app/opengraph-image.tsx supplies og:image and twitter:image.
// conventions. app/icon.png and app/apple-icon.png are the source.
//
// og:image and twitter:image are NOT inherited from
// app/opengraph-image.tsx — see the note on openGraph.images below. The
// icon conventions do reach these pages; the opengraph-image one does not.
metadataBase: new URL(SITE_URL),
openGraph: {
type: 'website',
title: 'Compare Schools Side by Side | schoolcompare',
description:
'Put five English schools on one screen — SATs, GCSE results, Ofsted grades, and how close you had to live to get a place.',
/*
* Declared, not inherited.
*
* app/opengraph-image.tsx is a metadata file convention, and it does
* attach to routes in the app root segment — _not-found gets an og:image
* from it. It does not reach the site's pages, which live in the
* (frontend) route group whose own layout.tsx is a root layout. Staging
* served og:title, og:description, og:url, og:site_name and og:type with
* no og:image at all, so every link pasted into a chat rendered bare.
*
* The file stays at the app root: /robots.txt and /icon.png depend on it
* being there, and moving it is what broke those before. This points at
* the route it generates instead. metadataBase makes it absolute.
*/
images: ['/opengraph-image'],
url: SITE_URL,
siteName: 'schoolcompare',
},
@@ -75,14 +94,34 @@ export const metadata: Metadata = {
title: 'Compare Schools Side by Side | schoolcompare',
description:
'Put five English schools on one screen — SATs, GCSE results, Ofsted grades, and how close you had to live to get a place.',
// The card is summary_large_image; claiming that and supplying no image
// is worse than claiming a summary card.
images: ['/opengraph-image'],
},
};
export default function RootLayout({
/*
* The footer's About and Blog links are flagged, which makes this the one
* place on the site that reads a flag on every route.
*
* 604800 is deliberate and load-bearing: it is the revalidate every SEO route
* here already declares. Next pins a route to the LOWEST revalidate among its
* fetches, so reading flags at the 300s default would drop the whole school
* and place corpus from a weekly cache to a 5-minute one — a large origin-load
* regression to hide two footer links.
*
* The cost is latency in one direction only. The pages themselves read the
* same flags at their own floors and flip within minutes; the footer links
* follow within a week. Turning a feature on early therefore shows the page
* before its footer link, which is harmless. Turning one off leaves a link to
* a 404 until the cache turns over, so a rollback that matters wants a purge.
*/
export default async function RootLayout({
children,
}: Readonly<{
children: React.ReactNode;
}>) {
const flags = await getFlags(604800);
return (
// The font variable classes must sit on <html>, not <body>. globals.css
// declares --font-display on :root as var(--font-manrope) and --font-ui as
@@ -126,7 +165,10 @@ export default function RootLayout({
{children}
</main>
<ComparisonToast />
<Footer />
<Footer
aboutEnabled={flags.about_page === true}
blogEnabled={flags.blog === true}
/>
</ComparisonProvider>
</body>
</html>
File renamed without changes.
File renamed without changes.
@@ -7,6 +7,8 @@
import { fetchSchoolDetails, fetchSchools, fetchNationalAverages } from '@/lib/api';
import { notFound, redirect } from 'next/navigation';
import { SchoolDetailShell } from '@/components/school/SchoolDetailShell';
import { NearbyPlaces } from '@/components/school/NearbyPlaces';
import { schoolBreadcrumbJsonLd, type SchoolPlace } from '@/lib/jsonld';
import { PrimarySchoolSections } from '@/components/school/PrimarySchoolSections';
import { SecondarySchoolSections } from '@/components/school/SecondarySchoolSections';
import {
@@ -149,6 +151,10 @@ export default async function SchoolPage({ params }: SchoolPageProps) {
}
const { school_info, yearly_data, absence_data, ofsted, census, admissions, admissions_history, admission_distance, deprivation, finance, destinations } = data;
// Absent on an older API build; the module and the trail both degrade to
// nothing rather than throwing, which is how this shipped without a
// lockstep deploy of the two images.
const places: SchoolPlace[] = data.places ?? [];
// Redirect bare URN to canonical slug URL
const canonicalSlug = schoolUrl(urn, school_info.school_name).replace('/school/', '');
@@ -185,10 +191,19 @@ export default async function SchoolPage({ params }: SchoolPageProps) {
const primaryNavItems = buildNavItems(primaryFlags, navInput);
const secondaryNavItems = buildSecondaryNavItems(secondaryFlags, navInput);
// Generate JSON-LD structured data for SEO
/*
* `School`, not `EducationalOrganization`.
*
* Both are valid, but EducationalOrganization is the parent type covering
* universities, training providers and nurseries alike. School is the
* specific one, and a type that says what the page is about is the whole
* point of declaring it. Google's own guidance treats the narrower type as
* the correct choice where it applies.
*/
const structuredData = {
'@context': 'https://schema.org',
'@type': 'EducationalOrganization',
'@graph': [{
'@type': 'School',
name: school_info.school_name,
identifier: school_info.urn.toString(),
...(school_info.address && {
@@ -210,6 +225,15 @@ export default async function SchoolPage({ params }: SchoolPageProps) {
...(school_info.school_type && {
additionalType: school_info.school_type,
}),
},
// The trail the page sits at the end of. School pages carried no
// breadcrumb at all, while every place page already emitted one.
schoolBreadcrumbJsonLd({
name: school_info.school_name,
url: `/school/${slug}`,
places,
}),
],
};
return (
@@ -264,6 +288,7 @@ export default async function SchoolPage({ params }: SchoolPageProps) {
/>
</SchoolDetailShell>
)}
<NearbyPlaces places={places} />
</>
);
}
@@ -0,0 +1,16 @@
import type { Metadata } from 'next';
import config from '@payload-config';
import { NotFoundPage, generatePageMetadata } from '@payloadcms/next/views';
import { importMap } from '../importMap.js';
type Args = {
params: Promise<{ segments: string[] }>;
searchParams: Promise<{ [key: string]: string | string[] }>;
};
export const generateMetadata = ({ params, searchParams }: Args): Promise<Metadata> =>
generatePageMetadata({ config, params, searchParams });
export default function NotFound({ params, searchParams }: Args) {
return NotFoundPage({ config, importMap, params, searchParams });
}
@@ -0,0 +1,16 @@
import type { Metadata } from 'next';
import config from '@payload-config';
import { RootPage, generatePageMetadata } from '@payloadcms/next/views';
import { importMap } from '../importMap.js';
type Args = {
params: Promise<{ segments: string[] }>;
searchParams: Promise<{ [key: string]: string | string[] }>;
};
export const generateMetadata = ({ params, searchParams }: Args): Promise<Metadata> =>
generatePageMetadata({ config, params, searchParams });
export default function Page({ params, searchParams }: Args) {
return RootPage({ config, importMap, params, searchParams });
}
@@ -0,0 +1,54 @@
import { RscEntryLexicalCell as RscEntryLexicalCell_44fe37237e0ebf4470c9990d8cb7b07e } from '@payloadcms/richtext-lexical/rsc'
import { RscEntryLexicalField as RscEntryLexicalField_44fe37237e0ebf4470c9990d8cb7b07e } from '@payloadcms/richtext-lexical/rsc'
import { LexicalDiffComponent as LexicalDiffComponent_44fe37237e0ebf4470c9990d8cb7b07e } from '@payloadcms/richtext-lexical/rsc'
import { BlocksFeatureClient as BlocksFeatureClient_e70f5e05f09f93e00b997edb1ef0c864 } from '@payloadcms/richtext-lexical/client'
import { BoldFeatureClient as BoldFeatureClient_e70f5e05f09f93e00b997edb1ef0c864 } from '@payloadcms/richtext-lexical/client'
import { ItalicFeatureClient as ItalicFeatureClient_e70f5e05f09f93e00b997edb1ef0c864 } from '@payloadcms/richtext-lexical/client'
import { UnderlineFeatureClient as UnderlineFeatureClient_e70f5e05f09f93e00b997edb1ef0c864 } from '@payloadcms/richtext-lexical/client'
import { StrikethroughFeatureClient as StrikethroughFeatureClient_e70f5e05f09f93e00b997edb1ef0c864 } from '@payloadcms/richtext-lexical/client'
import { SubscriptFeatureClient as SubscriptFeatureClient_e70f5e05f09f93e00b997edb1ef0c864 } from '@payloadcms/richtext-lexical/client'
import { SuperscriptFeatureClient as SuperscriptFeatureClient_e70f5e05f09f93e00b997edb1ef0c864 } from '@payloadcms/richtext-lexical/client'
import { InlineCodeFeatureClient as InlineCodeFeatureClient_e70f5e05f09f93e00b997edb1ef0c864 } from '@payloadcms/richtext-lexical/client'
import { ParagraphFeatureClient as ParagraphFeatureClient_e70f5e05f09f93e00b997edb1ef0c864 } from '@payloadcms/richtext-lexical/client'
import { HeadingFeatureClient as HeadingFeatureClient_e70f5e05f09f93e00b997edb1ef0c864 } from '@payloadcms/richtext-lexical/client'
import { AlignFeatureClient as AlignFeatureClient_e70f5e05f09f93e00b997edb1ef0c864 } from '@payloadcms/richtext-lexical/client'
import { IndentFeatureClient as IndentFeatureClient_e70f5e05f09f93e00b997edb1ef0c864 } from '@payloadcms/richtext-lexical/client'
import { UnorderedListFeatureClient as UnorderedListFeatureClient_e70f5e05f09f93e00b997edb1ef0c864 } from '@payloadcms/richtext-lexical/client'
import { OrderedListFeatureClient as OrderedListFeatureClient_e70f5e05f09f93e00b997edb1ef0c864 } from '@payloadcms/richtext-lexical/client'
import { ChecklistFeatureClient as ChecklistFeatureClient_e70f5e05f09f93e00b997edb1ef0c864 } from '@payloadcms/richtext-lexical/client'
import { LinkFeatureClient as LinkFeatureClient_e70f5e05f09f93e00b997edb1ef0c864 } from '@payloadcms/richtext-lexical/client'
import { RelationshipFeatureClient as RelationshipFeatureClient_e70f5e05f09f93e00b997edb1ef0c864 } from '@payloadcms/richtext-lexical/client'
import { BlockquoteFeatureClient as BlockquoteFeatureClient_e70f5e05f09f93e00b997edb1ef0c864 } from '@payloadcms/richtext-lexical/client'
import { UploadFeatureClient as UploadFeatureClient_e70f5e05f09f93e00b997edb1ef0c864 } from '@payloadcms/richtext-lexical/client'
import { HorizontalRuleFeatureClient as HorizontalRuleFeatureClient_e70f5e05f09f93e00b997edb1ef0c864 } from '@payloadcms/richtext-lexical/client'
import { InlineToolbarFeatureClient as InlineToolbarFeatureClient_e70f5e05f09f93e00b997edb1ef0c864 } from '@payloadcms/richtext-lexical/client'
import { CollectionCards as CollectionCards_f9c02e79a4aed9a3924487c0cd4cafb1 } from '@payloadcms/next/rsc'
/** @type import('payload').ImportMap */
export const importMap = {
"@payloadcms/richtext-lexical/rsc#RscEntryLexicalCell": RscEntryLexicalCell_44fe37237e0ebf4470c9990d8cb7b07e,
"@payloadcms/richtext-lexical/rsc#RscEntryLexicalField": RscEntryLexicalField_44fe37237e0ebf4470c9990d8cb7b07e,
"@payloadcms/richtext-lexical/rsc#LexicalDiffComponent": LexicalDiffComponent_44fe37237e0ebf4470c9990d8cb7b07e,
"@payloadcms/richtext-lexical/client#BlocksFeatureClient": BlocksFeatureClient_e70f5e05f09f93e00b997edb1ef0c864,
"@payloadcms/richtext-lexical/client#BoldFeatureClient": BoldFeatureClient_e70f5e05f09f93e00b997edb1ef0c864,
"@payloadcms/richtext-lexical/client#ItalicFeatureClient": ItalicFeatureClient_e70f5e05f09f93e00b997edb1ef0c864,
"@payloadcms/richtext-lexical/client#UnderlineFeatureClient": UnderlineFeatureClient_e70f5e05f09f93e00b997edb1ef0c864,
"@payloadcms/richtext-lexical/client#StrikethroughFeatureClient": StrikethroughFeatureClient_e70f5e05f09f93e00b997edb1ef0c864,
"@payloadcms/richtext-lexical/client#SubscriptFeatureClient": SubscriptFeatureClient_e70f5e05f09f93e00b997edb1ef0c864,
"@payloadcms/richtext-lexical/client#SuperscriptFeatureClient": SuperscriptFeatureClient_e70f5e05f09f93e00b997edb1ef0c864,
"@payloadcms/richtext-lexical/client#InlineCodeFeatureClient": InlineCodeFeatureClient_e70f5e05f09f93e00b997edb1ef0c864,
"@payloadcms/richtext-lexical/client#ParagraphFeatureClient": ParagraphFeatureClient_e70f5e05f09f93e00b997edb1ef0c864,
"@payloadcms/richtext-lexical/client#HeadingFeatureClient": HeadingFeatureClient_e70f5e05f09f93e00b997edb1ef0c864,
"@payloadcms/richtext-lexical/client#AlignFeatureClient": AlignFeatureClient_e70f5e05f09f93e00b997edb1ef0c864,
"@payloadcms/richtext-lexical/client#IndentFeatureClient": IndentFeatureClient_e70f5e05f09f93e00b997edb1ef0c864,
"@payloadcms/richtext-lexical/client#UnorderedListFeatureClient": UnorderedListFeatureClient_e70f5e05f09f93e00b997edb1ef0c864,
"@payloadcms/richtext-lexical/client#OrderedListFeatureClient": OrderedListFeatureClient_e70f5e05f09f93e00b997edb1ef0c864,
"@payloadcms/richtext-lexical/client#ChecklistFeatureClient": ChecklistFeatureClient_e70f5e05f09f93e00b997edb1ef0c864,
"@payloadcms/richtext-lexical/client#LinkFeatureClient": LinkFeatureClient_e70f5e05f09f93e00b997edb1ef0c864,
"@payloadcms/richtext-lexical/client#RelationshipFeatureClient": RelationshipFeatureClient_e70f5e05f09f93e00b997edb1ef0c864,
"@payloadcms/richtext-lexical/client#BlockquoteFeatureClient": BlockquoteFeatureClient_e70f5e05f09f93e00b997edb1ef0c864,
"@payloadcms/richtext-lexical/client#UploadFeatureClient": UploadFeatureClient_e70f5e05f09f93e00b997edb1ef0c864,
"@payloadcms/richtext-lexical/client#HorizontalRuleFeatureClient": HorizontalRuleFeatureClient_e70f5e05f09f93e00b997edb1ef0c864,
"@payloadcms/richtext-lexical/client#InlineToolbarFeatureClient": InlineToolbarFeatureClient_e70f5e05f09f93e00b997edb1ef0c864,
"@payloadcms/next/rsc#CollectionCards": CollectionCards_f9c02e79a4aed9a3924487c0cd4cafb1
}
@@ -0,0 +1,20 @@
/*
* Payload's REST API, mounted at /cms-api rather than /api.
* See lib/payloadRoutes.ts — /api is the FastAPI proxy's catch-all.
*/
import config from '@payload-config';
import {
REST_DELETE,
REST_GET,
REST_OPTIONS,
REST_PATCH,
REST_POST,
REST_PUT,
} from '@payloadcms/next/routes';
export const GET = REST_GET(config);
export const POST = REST_POST(config);
export const DELETE = REST_DELETE(config);
export const PATCH = REST_PATCH(config);
export const PUT = REST_PUT(config);
export const OPTIONS = REST_OPTIONS(config);
@@ -0,0 +1,4 @@
import config from '@payload-config';
import { GRAPHQL_PLAYGROUND_GET } from '@payloadcms/next/routes';
export const GET = GRAPHQL_PLAYGROUND_GET(config);
@@ -0,0 +1,5 @@
import config from '@payload-config';
import { GRAPHQL_POST, REST_OPTIONS } from '@payloadcms/next/routes';
export const POST = GRAPHQL_POST(config);
export const OPTIONS = REST_OPTIONS(config);
+27
View File
@@ -0,0 +1,27 @@
/**
* Root layout for the Payload admin panel.
*
* This is a SECOND root layout: it renders its own <html>/<body>, as does
* app/(frontend)/layout.tsx. Next permits that only while no app/layout.tsx
* exists — which is why the site's routes were moved into (frontend). Adding
* an app/layout.tsx would nest the admin panel inside the site's nav, footer
* and providers and emit nested <html>.
*/
import type { ServerFunctionClient } from 'payload';
import config from '@payload-config';
import { RootLayout, handleServerFunctions } from '@payloadcms/next/layouts';
import { importMap } from './admin/importMap.js';
import '@payloadcms/next/css';
const serverFunction: ServerFunctionClient = async function (args) {
'use server';
return handleServerFunctions({ ...args, config, importMap });
};
export default function PayloadLayout({ children }: { children: React.ReactNode }) {
return (
<RootLayout config={config} importMap={importMap} serverFunction={serverFunction}>
{children}
</RootLayout>
);
}
+7 -2
View File
@@ -12,9 +12,14 @@ export default function robots(): MetadataRoute.Robots {
{
userAgent: '*',
allow: '/',
disallow: ['/api/', '/_next/'],
// /admin and /cms-api are also served X-Robots-Tag: noindex by
// next.config.mjs. A Disallow alone blocks crawling, not indexing.
disallow: ['/api/', '/_next/', '/admin/', '/cms-api/'],
},
],
sitemap: absoluteUrl('/sitemap.xml'),
// Two sitemaps: /sitemap.xml is proxied from FastAPI and carries the
// school corpus; /content-sitemap.xml is Next-owned and carries /about
// and the blog. The backend knows nothing about Payload.
sitemap: [absoluteUrl('/sitemap.xml'), absoluteUrl('/content-sitemap.xml')],
};
}
+25
View File
@@ -0,0 +1,25 @@
import type { Block } from 'payload';
/**
* The house block: "what this number doesn't tell you".
*
* Blocks are the reason this site runs a CMS rather than flat files — a post
* can carry live product components, not screenshots of them. This is the
* first and simplest one; a live-chart block follows when a post needs it.
*/
export const Callout: Block = {
slug: 'callout',
labels: { singular: 'Callout', plural: 'Callouts' },
fields: [
{
name: 'tone',
type: 'select',
defaultValue: 'caveat',
options: [
{ label: 'Caveat: what this does not show', value: 'caveat' },
{ label: 'Note: useful aside', value: 'note' },
],
},
{ name: 'body', type: 'textarea', required: true },
],
};
+31
View File
@@ -0,0 +1,31 @@
import type { CollectionConfig } from 'payload';
/**
* Uploads land on a Docker named volume mounted at /app/media. The path is
* absolute because Payload 3 requires it, and it must match the payload_media
* mount in docker-compose.portainer.yml exactly — a mismatch writes into the
* container's own filesystem, where the next redeploy silently discards it.
*/
export const Media: CollectionConfig = {
slug: 'media',
access: { read: () => true },
upload: {
staticDir: '/app/media',
mimeTypes: ['image/*'],
imageSizes: [
{ name: 'thumbnail', width: 400 },
{ name: 'hero', width: 1200 },
],
adminThumbnail: 'thumbnail',
},
fields: [
{
name: 'alt',
type: 'text',
required: true,
// Required rather than optional: a decorative-by-default image is an
// accessibility regression on a site parents use under time pressure.
admin: { description: 'Describe the image for screen readers.' },
},
],
};
+99
View File
@@ -0,0 +1,99 @@
import type { CollectionConfig } from 'payload';
import { revalidatePath } from 'next/cache';
import { lexicalEditor, BlocksFeature } from '@payloadcms/richtext-lexical';
import { Callout } from '@/blocks/Callout';
/**
* Drop the cached copy of a post page when it changes.
*
* Only the post page needs this. The blog index, the RSS feed and the content
* sitemap are force-dynamic — they have no dynamic params, so Next would
* prerender them at build time, where CI has no database — which means they
* already reflect a change on the next request.
*
* /blog/[slug] is ISR: generated on first request and cached, so without this
* an edit to an already-published post would not appear until the revalidate
* window expired — up to an hour of a writer concluding that saving is broken.
*
* Payload runs in the same process as Next, so this is a direct revalidatePath
* call: no webhook, no shared secret, no network hop to get wrong.
*/
function revalidatePost(slug: string) {
revalidatePath(`/blog/${slug}`);
}
export const Posts: CollectionConfig = {
slug: 'posts',
access: {
/*
* Drafts must be hidden here, not in the pages that query this collection.
*
* From Payload's own documentation: "The `draft` argument alone does not
* restrict documents with `_status: 'draft'` from being returned by the
* API." The blog index and post page both filter on `_status`, but that
* is a convenience, not a control — a direct GET /cms-api/posts would
* hand every unpublished draft to any visitor.
*
* Returning a query constraint rather than a boolean is the documented
* mechanism: Payload merges it into every read for an anonymous caller.
*/
read: ({ req }) => {
if (req.user) return true;
return { _status: { equals: 'published' } };
},
},
admin: {
useAsTitle: 'title',
defaultColumns: ['title', 'publishedAt', '_status'],
},
versions: {
// Posts get written across several sittings and previewed before they go
// live. Without drafts, saving is publishing.
drafts: true,
},
hooks: {
afterChange: [({ doc }) => { revalidatePost(String(doc.slug)); }],
afterDelete: [({ doc }) => { revalidatePost(String(doc.slug)); }],
},
fields: [
{ name: 'title', type: 'text', required: true },
{
name: 'slug',
type: 'text',
required: true,
unique: true,
index: true,
admin: {
position: 'sidebar',
description: 'The URL segment. Never change it after publishing.',
},
},
{
name: 'publishedAt',
type: 'date',
required: true,
admin: { position: 'sidebar', date: { pickerAppearance: 'dayOnly' } },
},
{
name: 'excerpt',
type: 'textarea',
required: true,
maxLength: 200,
admin: {
description: 'Shown on the index and used as the meta description.',
},
},
{ name: 'heroImage', type: 'upload', relationTo: 'media' },
{
name: 'content',
type: 'richText',
required: true,
editor: lexicalEditor({
features: ({ defaultFeatures }) => [
...defaultFeatures,
BlocksFeature({ blocks: [Callout] }),
],
}),
},
],
};
+33
View File
@@ -0,0 +1,33 @@
import type { CollectionConfig } from 'payload';
/**
* The site's only authenticated surface. There is one account and no
* registration: `create` is closed to everyone, so the first user is seeded
* with `payload create-first-user` and no one can add another through the API.
*/
export const Users: CollectionConfig = {
slug: 'users',
auth: {
// Slows credential stuffing against a panel that is on the public
// internet. Five attempts, then a ten-minute lock.
maxLoginAttempts: 5,
lockTime: 10 * 60 * 1000,
},
access: {
create: () => false,
read: ({ req }) => Boolean(req.user),
update: ({ req }) => Boolean(req.user),
delete: () => false,
},
admin: { useAsTitle: 'email' },
fields: [
{
name: 'displayName',
type: 'text',
required: true,
// Rendered as the byline on every post. First name only — the site
// publishes no surname and no employer.
defaultValue: 'Tudor',
},
],
};
+10 -1
View File
@@ -22,7 +22,8 @@
.content {
display: grid;
grid-template-columns: 1.6fr 1fr 1fr;
/* Brand column plus three link columns: Product, Resources, About. */
grid-template-columns: 1.6fr 1fr 1fr 1fr;
gap: 2rem;
margin-bottom: 3rem;
}
@@ -193,6 +194,14 @@
color: var(--on-sunken);
}
/* Four columns crush between the tablet range and the 768px collapse, so
pair them up first rather than jumping straight to a single column. */
@media (max-width: 960px) {
.content {
grid-template-columns: 1fr 1fr;
}
}
@media (max-width: 768px) {
.container {
padding: 2rem 1rem 1.5rem;
+32 -1
View File
@@ -10,7 +10,17 @@
import { LogoMark } from './Logo';
import styles from './Footer.module.css';
export function Footer() {
/**
* Both default to false so a caller that forgets a prop hides the link rather
* than pointing it at a page that 404s. Same reasoning as backend/flags.py:
* "Every flag defaults to False."
*/
interface FooterProps {
aboutEnabled?: boolean;
blogEnabled?: boolean;
}
export function Footer({ aboutEnabled = false, blogEnabled = false }: FooterProps = {}) {
const currentYear = new Date().getFullYear();
return (
@@ -93,6 +103,27 @@ export function Footer() {
</li>
</ul>
</div>
{/* Dropped entirely when both flags are dark, rather than left as an
empty heading: shipping dark means the footer renders as it did
before the feature existed. */}
{(aboutEnabled || blogEnabled) && (
<div className={styles.section}>
<h4 className={styles.sectionTitle}>About</h4>
<ul className={styles.links}>
{/* The only route to a named human. Deliberately not in the nav:
the mobile bottom bar already carries four items, and both of
these are lower intent than any of them. Post bylines link
here too, which is where a reader actually asks the question. */}
{aboutEnabled && (
<li><a href="/about" className={styles.link}>Who&apos;s behind this</a></li>
)}
{blogEnabled && (
<li><a href="/blog" className={styles.link}>Blog</a></li>
)}
</ul>
</div>
)}
</div>
<div className={styles.bottom}>
@@ -0,0 +1,21 @@
.caveat,
.note {
border-left: 3px solid var(--brand);
background: var(--brand-bg);
padding: 1rem 1.15rem;
margin: 1.75rem 0;
border-radius: 0 8px 8px 0;
}
.note {
border-left-color: var(--border-strong);
background: var(--bg-secondary);
}
.body {
font-family: var(--font-ui);
font-size: 0.95rem;
line-height: 1.65;
color: var(--text-primary);
margin: 0;
}
@@ -0,0 +1,14 @@
import styles from './CalloutBlock.module.css';
/**
* Renders the Callout block from blocks/Callout.ts. The "caveat" tone is the
* one that matters: it is how a post says what a number does not show, in
* context, rather than burying it in a closing paragraph.
*/
export function CalloutBlock({ tone, body }: { tone: string; body: string }) {
return (
<aside className={tone === 'caveat' ? styles.caveat : styles.note}>
<p className={styles.body}>{body}</p>
</aside>
);
}
@@ -0,0 +1,53 @@
/* Tokens only — the same vocabulary schoolSections.module.css uses, so the
module follows both themes without a rule of its own. No hardcoded colour
appears here; darkThemeSafety asserts that across the codebase. */
.section {
margin-top: 2rem;
}
/* Matches .sectionTitle in schoolSections.module.css, including the brand
rule before the text, so this reads as one more section of the page
rather than a footer bolted underneath it. */
.heading {
font-size: 1.125rem;
font-weight: 600;
color: var(--text-primary);
margin-bottom: 0.875rem;
padding-bottom: 0.5rem;
border-bottom: 2px solid var(--border);
font-family: var(--font-display);
display: flex;
align-items: center;
gap: 0.375rem;
}
.heading::before {
content: "";
display: inline-block;
width: 3px;
height: 1em;
background: var(--brand);
border-radius: 2px;
flex-shrink: 0;
}
.list {
display: flex;
flex-wrap: wrap;
gap: 0.5rem 1.25rem;
list-style: none;
margin: 0;
padding: 0;
}
.link {
color: var(--brand-strong);
font-weight: 500;
text-decoration: underline;
text-underline-offset: 2px;
}
.link:hover {
text-decoration-thickness: 2px;
}
@@ -0,0 +1,67 @@
import Link from 'next/link';
import type { SchoolPlace, SchoolPhasePage } from '@/lib/jsonld';
import styles from './NearbyPlaces.module.css';
/**
* Links from a school page into the location layer.
*
* This exists for a structural reason rather than a decorative one. Before
* it, the only anchor on a school page pointed at the school's own website,
* so the ~27k pages that carry most of the site's inbound authority passed it
* straight off-site and none of it reached the place pages. These links are
* what circulate it instead.
*
* Every entry comes from the place registry via the API, so a link is only
* ever offered for a page that exists: a place below the publish threshold is
* absent from the registry and therefore absent here.
*/
/** Narrowest first: a reader on a school page wants its town before its
* county. The API orders widest-first because that is what the breadcrumb
* reads, so the two orders are deliberately different. */
const ORDER: Record<string, number> = {
town: 0, locality: 0, outcode: 1, authority: 2,
};
function label(place: SchoolPlace): string {
const noun = place.count === 1 ? 'school' : 'schools';
const preposition = place.kind === 'outcode' ? 'near' : 'in';
return `${place.count} ${noun} ${preposition} ${place.name}`;
}
/** "22 primary schools in Brentwood" — the phrasing the query itself uses. */
function phaseLabel(place: SchoolPlace, page: SchoolPhasePage): string {
const noun = page.count === 1 ? 'school' : 'schools';
return `${page.count} ${page.phase} ${noun} in ${place.name}`;
}
export function NearbyPlaces({ places }: { places: SchoolPlace[] }) {
if (places.length === 0) return null;
const sorted = [...places].sort(
(a, b) => (ORDER[a.kind] ?? 9) - (ORDER[b.kind] ?? 9),
);
return (
<section className={styles.section} aria-labelledby="nearby-places">
<h2 id="nearby-places" className={styles.heading}>More schools near here</h2>
<ul className={styles.list}>
{sorted.flatMap((place) => [
<li key={`${place.kind}:${place.slug}`}>
<Link href={place.url} className={styles.link}>{label(place)}</Link>
</li>,
/* Immediately after its own place, so "22 primary schools in
Brentwood" reads as part of Brentwood rather than as an
unrelated link further down the row. */
...place.phases.map((page) => (
<li key={`${place.kind}:${place.slug}:${page.phase}`}>
<Link href={page.url} className={styles.link}>
{phaseLabel(place, page)}
</Link>
</li>
)),
])}
</ul>
</section>
);
}
+106
View File
@@ -0,0 +1,106 @@
# Publishing to the blog
The blog is Payload CMS, running inside the Next.js app. There is no separate
service and no second deploy. Writing a post is done in the browser and takes
effect on the live site within seconds.
## Before any of this: the flags
`/blog` and `/about` are behind feature flags (`blog` and `about_page`), and
every flag in this system starts off. While `blog` is dark, `/blog`, every post
page and the RSS feed return 404, and neither appears in the sitemap or the
footer. Posts still save normally, because `/admin` is deliberately **not**
flagged: you have to be able to write a post before there is anything worth
switching on.
So a new post published to a dark blog is invisible, and that is working as
intended, not a bug. Flip the flag in Unleash when the content is ready.
Staging and production hold their own values (`development` and `production`
environments), so you can light it on staging first. The page follows a flip
within five minutes; the footer link takes up to a week, because it renders in
the root layout and is cached at the same weekly floor as the school corpus.
Both flags are temporary scaffolding, like every flag here: a test starts
failing once one is older than 90 days, at which point either the feature is
permanent and the flag comes out, or it was never going to ship.
## Signing in
`https://www.schoolcompare.co.uk/admin`, one account, no registration. If you
need the account seeded on a fresh environment, run against the container:
```bash
npx payload create-first-user
```
Staging has its own admin panel, its own database and its own credentials at
`https://stx.schoolcompare.co.uk/admin`. Never reuse production's secret or
password there.
## Writing a post
**Posts → Create New.** The fields:
| Field | Notes |
|---|---|
| **Title** | The `<h1>` and the browser tab. |
| **Slug** | The URL segment, in the sidebar. **Never change it after publishing.** It is the canonical URL, and changing it breaks every existing link and discards the page's accumulated search signal. |
| **Published at** | The date shown on the post and in the feed. |
| **Excerpt** | Max 200 characters. Shown on the index *and* used as the meta description, so write it as a standalone sentence rather than a teaser. |
| **Hero image** | Optional. Becomes the social share image; without one, the site's generated card is used. |
| **Content** | Rich text. `/` inserts a block. |
**Save as draft** while you're working; drafts are not public. **Publish** when
it's ready.
### The callout block
One custom block, `Callout`, with two tones:
- **Caveat**: what a number does *not* show. This is the one that matters. It
is how a post states a limitation in context rather than burying it in a
closing paragraph.
- **Note**: a useful aside.
### Images
Every image requires alt text; the editor will not let you save without it.
Uploads go to a Docker volume on the host, which is backed up separately from
Postgres. An image is not reproducible from the pipeline the way school data
is.
## How publishing reaches the live site
- `/blog`, `/blog/rss.xml` and `/content-sitemap.xml` are rendered per request,
so a new post appears immediately.
- `/blog/[slug]` is cached after its first request. Publishing or editing fires
a `revalidatePath` from the collection's `afterChange` hook, which drops that
cached copy, so edits appear immediately too.
If a change doesn't show, it is far more likely the post is still a draft than
that the cache is stale.
## House style
These rules are why the blog exists. A post that ignores them makes the site
read more machine-generated, not less.
- **First person singular.** "I built", "I found", never "we provide".
- **Concrete over general.** "When we were looking at schools in Wandsworth"
beats any amount of stated warmth.
- **State limits before someone else finds them.** Every post that presents a
metric says what it does not show. This is the single strongest signal that a
human wrote it: generated content does not volunteer its own weaknesses.
- **No mission statements, no "passionate about", no invented team.** There is
one person here.
- **No em dashes.** They are one of the clearest tells of machine-written
prose, which is the whole problem this blog exists to fix. A full stop, a
colon, a semicolon or a pair of commas does the job and reads as though a
person chose it.
- **Short sentences.**
- **Never publish a surname, an employer, or a child's name.** The site's author
is "Tudor". See `/about`.
- **Never invent a figure**, even illustratively. On a site whose whole
proposition is official data, a made-up number attached to a real school is
the one thing it cannot do, and no illustrative intent survives being
screenshotted.
File renamed without changes.
-27
View File
@@ -311,26 +311,6 @@ export async function fetchDataInfo(
return handleResponse<DataInfoResponse>(response);
}
// ============================================================================
// Client-Side Fetcher (for SWR)
// ============================================================================
/**
* Generic fetcher function for use with SWR
* @example
* ```tsx
* const { data, error } = useSWR('/api/schools', fetcher);
* ```
*/
export async function fetcher<T>(url: string): Promise<T> {
// If it's already a full URL, use it directly
// Otherwise, prepend the API_BASE_URL
const fullUrl = url.startsWith('http') ? url : `${API_BASE_URL}${url.startsWith('/') ? url : `/${url}`}`;
const response = await fetch(fullUrl);
return handleResponse<T>(response);
}
// ============================================================================
// Geocoding API
// ============================================================================
@@ -396,10 +376,3 @@ export function calculateDistance(
const c = 2 * Math.atan2(Math.sqrt(a), Math.sqrt(1 - a));
return R * c;
}
/**
* Convert kilometers to miles
*/
export function kmToMiles(km: number): number {
return km * 0.621371;
}
+13 -3
View File
@@ -26,11 +26,21 @@ export const FLAGS_REVALIDATE = 300;
const API = process.env.FASTAPI_URL || process.env.NEXT_PUBLIC_API_URL
|| 'http://localhost:8000/api';
/** Every flag and its value. Never throws: an unreadable flag is a dark one. */
export async function getFlags(): Promise<Flags> {
/**
* Every flag and its value. Never throws: an unreadable flag is a dark one.
*
* `revalidate` is the caller's, because reading flags pins the whole route to
* the lowest revalidate among its fetches. Every SEO route here declares
* 604800; gating one at the 300s default would drop it from a weekly cache to
* a 5-minute one. Pass the route's own floor and gating costs it nothing.
*
* The trade is flag-flip latency: a route that revalidates weekly takes up to
* a week to notice a flip. Pass a smaller number where a flip must land fast.
*/
export async function getFlags(revalidate: number = FLAGS_REVALIDATE): Promise<Flags> {
try {
const res = await fetch(`${API}/flags`, {
next: { revalidate: FLAGS_REVALIDATE },
next: { revalidate },
});
if (!res.ok) return {};
return await res.json();
+159
View File
@@ -0,0 +1,159 @@
import { SITE_URL, absoluteUrl } from '@/lib/site';
/**
* The site's author entity.
*
* First name only, by choice — see /about. That makes this a weaker search
* signal than a fully identified author would be, which is why the About page
* carries a substantial methodology section: the credibility has to come from
* stated provenance rather than from a corroborable identity.
*
* Everything that needs an author — the About page, every post byline —
* references this one shape, so search engines resolve them all to one entity.
*/
export function personJsonLd() {
return {
'@type': 'Person',
'@id': `${SITE_URL}/about#tudor`,
name: 'Tudor',
url: absoluteUrl('/about'),
image: absoluteUrl('/brand/tudor.jpg'),
description:
'Parent in south-west London who built schoolcompare while looking for a primary school.',
} as const;
}
export function organizationJsonLd() {
return {
'@type': 'Organization',
'@id': `${SITE_URL}#organization`,
name: 'schoolcompare',
url: SITE_URL,
logo: absoluteUrl('/icon-512.png'),
} as const;
}
interface PostSummary {
title: string;
slug: string;
excerpt: string;
publishedAt: string;
}
/**
* References the Person and Organization by @id rather than repeating them, so
* search engines resolve every post and the About page to the one author
* entity. Repeating the shape would declare several people with one name.
*
* `namedAuthor` is the about_page flag. The Person entity is anchored at
* /about#tudor, and that URL 404s while the flag is dark, so a post published
* in that state must not claim it: an author @id resolving to nothing is a
* worse signal than no named author. It falls back to the publisher, which is
* always live. The parameter is required rather than defaulted because every
* call site has the flag to hand and the wrong default is silent.
*/
export function blogPostingJsonLd(
post: PostSummary,
{ namedAuthor }: { namedAuthor: boolean },
) {
return {
'@type': 'BlogPosting',
headline: post.title,
description: post.excerpt,
url: absoluteUrl(`/blog/${post.slug}`),
datePublished: post.publishedAt,
author: {
'@id': namedAuthor ? `${SITE_URL}/about#tudor` : `${SITE_URL}#organization`,
},
publisher: { '@id': `${SITE_URL}#organization` },
} as const;
}
/**
* A place the location layer publishes a page for, as the school API reports
* it. `count` is what lets a link say "All 37 schools in Brentwood" rather
* than "click here".
*/
export interface SchoolPhasePage {
phase: string;
count: number;
url: string;
}
export interface SchoolPlace {
kind: string;
slug: string;
name: string;
count: number;
url: string;
/**
* The phase variants this school is actually listed on: usually one, two
* for an all-through school, none for an outcode, which publishes no phase
* route. Decided by the place registry, never re-derived here.
*/
phases: SchoolPhasePage[];
}
/**
* The trail a school page sits at the end of: Schools → authority → town.
*
* Only authority and town/locality appear. An outcode is a useful link in the
* module beside this — a parent does search "schools near CM15" — but it is
* not a step anyone navigates through, and a breadcrumb that claims otherwise
* describes a hierarchy the site does not have.
*
* Levels are skipped rather than faked. A school whose town falls below the
* publish threshold has no town page, so the trail closes over the gap; the
* alternative is a breadcrumb linking to a 404.
*/
export function schoolBreadcrumbJsonLd(
school: { name: string; url: string; places: SchoolPlace[] },
) {
/*
* Rooted at the homepage, not at /schools. There is no /schools index page
* — the location layer is /schools/[place], /schools/authority/[la] and
* /schools/near/[outcode], with nothing at the bare path — so a trail
* starting there would open with a link to a 404.
*/
const trail: Array<{ name: string; url: string }> = [
{ name: 'schoolcompare', url: '/' },
];
const authority = school.places.find((p) => p.kind === 'authority');
if (authority) trail.push({ name: authority.name, url: authority.url });
const town = school.places.find((p) => p.kind === 'town' || p.kind === 'locality');
if (town) trail.push({ name: town.name, url: town.url });
trail.push({ name: school.name, url: school.url });
return {
'@type': 'BreadcrumbList',
itemListElement: trail.map((step, index) => ({
'@type': 'ListItem',
position: index + 1,
name: step.name,
item: absoluteUrl(step.url),
})),
} as const;
}
export function breadcrumbJsonLd(post: PostSummary) {
return {
'@type': 'BreadcrumbList',
itemListElement: [
{
'@type': 'ListItem',
position: 1,
name: 'Blog',
item: absoluteUrl('/blog'),
},
{
'@type': 'ListItem',
position: 2,
name: post.title,
item: absoluteUrl(`/blog/${post.slug}`),
},
],
} as const;
}
+15
View File
@@ -0,0 +1,15 @@
import { getPayload } from 'payload';
import config from '@payload-config';
import type { Payload } from 'payload';
/**
* One Payload instance per process. getPayload() is itself memoised by
* Payload, but routing every caller through here keeps the config import in a
* single place and gives page code one name to mock in tests.
*
* Never call this at module scope: CI builds the image with no database
* reachable, so a build-time connection attempt fails the build.
*/
export function getCachedPayload(): Promise<Payload> {
return getPayload({ config });
}
+23
View File
@@ -0,0 +1,23 @@
/**
* Where Payload mounts, defined once.
*
* These are imported by payload.config.ts and asserted by
* __tests__/payload/routes.test.ts. They live in their own module because
* payload.config.ts cannot be imported from a Jest test: Payload ships
* ESM-only, and next/jest's transformIgnorePatterns skips node_modules — you
* cannot un-ignore a package by appending patterns, and forcing it through
* `transpilePackages` would change how the production build bundles Payload
* to serve a test. Keeping the values here makes them testable without
* loading Payload at all.
*/
/**
* Payload's API base. It must NOT be '/api': that path belongs to
* app/(frontend)/api/[...path]/route.ts, a catch-all that proxies to FastAPI.
* It would swallow every admin API call and forward it to the backend, and
* the failure is silent — no error, just wrong responses.
*/
export const PAYLOAD_API_ROUTE = '/cms-api';
/** The admin panel. Kept out of the index by robots.txt and X-Robots-Tag. */
export const PAYLOAD_ADMIN_ROUTE = '/admin';
+11
View File
@@ -1,3 +1,5 @@
import type { SchoolPlace } from '@/lib/jsonld';
/**
* TypeScript type definitions for SchoolCompare API
* Generated from backend/models.py and backend/schemas.py
@@ -346,6 +348,15 @@ export interface SchoolsResponse {
export interface SchoolDetailsResponse {
school_info: School;
/**
* The published location-layer pages containing this school, widest first.
*
* Optional because the frontend and backend ship as separate images: a
* frontend deployed ahead of the API that serves this must render without
* it, not throw. Empty is also a real answer — a school whose town and
* authority both fall below the publish threshold has nowhere to link.
*/
places?: SchoolPlace[];
yearly_data: SchoolResult[];
absence_data: AbsenceData | null;
// Supplementary data (null until Kestra populates)
File diff suppressed because it is too large. Load diff
@@ -0,0 +1,216 @@
import { MigrateUpArgs, MigrateDownArgs, sql } from '@payloadcms/db-postgres'
export async function up({ db, payload, req }: MigrateUpArgs): Promise<void> {
/*
* Hand-added, and it must survive any regeneration of this file.
*
* `schemaName: 'payload'` tells Payload where to put its tables; it does not
* create the schema. Every statement below is qualified to "payload", so on
* a database that has never run this (staging and production both), the
* whole migration fails with `schema "payload" does not exist`. The schema
* only existed on the throwaway database used to generate this because it
* was created there by hand.
*/
await db.execute(sql`CREATE SCHEMA IF NOT EXISTS "payload";`)
await db.execute(sql`
CREATE TYPE "payload"."enum_posts_status" AS ENUM('draft', 'published');
CREATE TYPE "payload"."enum__posts_v_version_status" AS ENUM('draft', 'published');
CREATE TABLE "payload"."users_sessions" (
"_order" integer NOT NULL,
"_parent_id" integer NOT NULL,
"id" varchar PRIMARY KEY NOT NULL,
"created_at" timestamp(3) with time zone,
"expires_at" timestamp(3) with time zone NOT NULL
);
CREATE TABLE "payload"."users" (
"id" serial PRIMARY KEY NOT NULL,
"display_name" varchar DEFAULT 'Tudor' NOT NULL,
"updated_at" timestamp(3) with time zone DEFAULT now() NOT NULL,
"created_at" timestamp(3) with time zone DEFAULT now() NOT NULL,
"email" varchar NOT NULL,
"reset_password_token" varchar,
"reset_password_expiration" timestamp(3) with time zone,
"salt" varchar,
"hash" varchar,
"login_attempts" numeric DEFAULT 0,
"lock_until" timestamp(3) with time zone
);
CREATE TABLE "payload"."posts" (
"id" serial PRIMARY KEY NOT NULL,
"title" varchar,
"slug" varchar,
"published_at" timestamp(3) with time zone,
"excerpt" varchar,
"hero_image_id" integer,
"content" jsonb,
"updated_at" timestamp(3) with time zone DEFAULT now() NOT NULL,
"created_at" timestamp(3) with time zone DEFAULT now() NOT NULL,
"_status" "payload"."enum_posts_status" DEFAULT 'draft'
);
CREATE TABLE "payload"."_posts_v" (
"id" serial PRIMARY KEY NOT NULL,
"parent_id" integer,
"version_title" varchar,
"version_slug" varchar,
"version_published_at" timestamp(3) with time zone,
"version_excerpt" varchar,
"version_hero_image_id" integer,
"version_content" jsonb,
"version_updated_at" timestamp(3) with time zone,
"version_created_at" timestamp(3) with time zone,
"version__status" "payload"."enum__posts_v_version_status" DEFAULT 'draft',
"created_at" timestamp(3) with time zone DEFAULT now() NOT NULL,
"updated_at" timestamp(3) with time zone DEFAULT now() NOT NULL,
"latest" boolean
);
CREATE TABLE "payload"."media" (
"id" serial PRIMARY KEY NOT NULL,
"alt" varchar NOT NULL,
"updated_at" timestamp(3) with time zone DEFAULT now() NOT NULL,
"created_at" timestamp(3) with time zone DEFAULT now() NOT NULL,
"url" varchar,
"thumbnail_u_r_l" varchar,
"filename" varchar,
"mime_type" varchar,
"filesize" numeric,
"width" numeric,
"height" numeric,
"focal_x" numeric,
"focal_y" numeric,
"sizes_thumbnail_url" varchar,
"sizes_thumbnail_width" numeric,
"sizes_thumbnail_height" numeric,
"sizes_thumbnail_mime_type" varchar,
"sizes_thumbnail_filesize" numeric,
"sizes_thumbnail_filename" varchar,
"sizes_hero_url" varchar,
"sizes_hero_width" numeric,
"sizes_hero_height" numeric,
"sizes_hero_mime_type" varchar,
"sizes_hero_filesize" numeric,
"sizes_hero_filename" varchar
);
CREATE TABLE "payload"."payload_kv" (
"id" serial PRIMARY KEY NOT NULL,
"key" varchar NOT NULL,
"data" jsonb NOT NULL
);
CREATE TABLE "payload"."payload_locked_documents" (
"id" serial PRIMARY KEY NOT NULL,
"global_slug" varchar,
"updated_at" timestamp(3) with time zone DEFAULT now() NOT NULL,
"created_at" timestamp(3) with time zone DEFAULT now() NOT NULL
);
CREATE TABLE "payload"."payload_locked_documents_rels" (
"id" serial PRIMARY KEY NOT NULL,
"order" integer,
"parent_id" integer NOT NULL,
"path" varchar NOT NULL,
"users_id" integer,
"posts_id" integer,
"media_id" integer
);
CREATE TABLE "payload"."payload_preferences" (
"id" serial PRIMARY KEY NOT NULL,
"key" varchar,
"value" jsonb,
"updated_at" timestamp(3) with time zone DEFAULT now() NOT NULL,
"created_at" timestamp(3) with time zone DEFAULT now() NOT NULL
);
CREATE TABLE "payload"."payload_preferences_rels" (
"id" serial PRIMARY KEY NOT NULL,
"order" integer,
"parent_id" integer NOT NULL,
"path" varchar NOT NULL,
"users_id" integer
);
CREATE TABLE "payload"."payload_migrations" (
"id" serial PRIMARY KEY NOT NULL,
"name" varchar,
"batch" numeric,
"updated_at" timestamp(3) with time zone DEFAULT now() NOT NULL,
"created_at" timestamp(3) with time zone DEFAULT now() NOT NULL
);
ALTER TABLE "payload"."users_sessions" ADD CONSTRAINT "users_sessions_parent_id_fk" FOREIGN KEY ("_parent_id") REFERENCES "payload"."users"("id") ON DELETE cascade ON UPDATE no action;
ALTER TABLE "payload"."posts" ADD CONSTRAINT "posts_hero_image_id_media_id_fk" FOREIGN KEY ("hero_image_id") REFERENCES "payload"."media"("id") ON DELETE set null ON UPDATE no action;
ALTER TABLE "payload"."_posts_v" ADD CONSTRAINT "_posts_v_parent_id_posts_id_fk" FOREIGN KEY ("parent_id") REFERENCES "payload"."posts"("id") ON DELETE set null ON UPDATE no action;
ALTER TABLE "payload"."_posts_v" ADD CONSTRAINT "_posts_v_version_hero_image_id_media_id_fk" FOREIGN KEY ("version_hero_image_id") REFERENCES "payload"."media"("id") ON DELETE set null ON UPDATE no action;
ALTER TABLE "payload"."payload_locked_documents_rels" ADD CONSTRAINT "payload_locked_documents_rels_parent_fk" FOREIGN KEY ("parent_id") REFERENCES "payload"."payload_locked_documents"("id") ON DELETE cascade ON UPDATE no action;
ALTER TABLE "payload"."payload_locked_documents_rels" ADD CONSTRAINT "payload_locked_documents_rels_users_fk" FOREIGN KEY ("users_id") REFERENCES "payload"."users"("id") ON DELETE cascade ON UPDATE no action;
ALTER TABLE "payload"."payload_locked_documents_rels" ADD CONSTRAINT "payload_locked_documents_rels_posts_fk" FOREIGN KEY ("posts_id") REFERENCES "payload"."posts"("id") ON DELETE cascade ON UPDATE no action;
ALTER TABLE "payload"."payload_locked_documents_rels" ADD CONSTRAINT "payload_locked_documents_rels_media_fk" FOREIGN KEY ("media_id") REFERENCES "payload"."media"("id") ON DELETE cascade ON UPDATE no action;
ALTER TABLE "payload"."payload_preferences_rels" ADD CONSTRAINT "payload_preferences_rels_parent_fk" FOREIGN KEY ("parent_id") REFERENCES "payload"."payload_preferences"("id") ON DELETE cascade ON UPDATE no action;
ALTER TABLE "payload"."payload_preferences_rels" ADD CONSTRAINT "payload_preferences_rels_users_fk" FOREIGN KEY ("users_id") REFERENCES "payload"."users"("id") ON DELETE cascade ON UPDATE no action;
CREATE INDEX "users_sessions_order_idx" ON "payload"."users_sessions" USING btree ("_order");
CREATE INDEX "users_sessions_parent_id_idx" ON "payload"."users_sessions" USING btree ("_parent_id");
CREATE INDEX "users_updated_at_idx" ON "payload"."users" USING btree ("updated_at");
CREATE INDEX "users_created_at_idx" ON "payload"."users" USING btree ("created_at");
CREATE UNIQUE INDEX "users_email_idx" ON "payload"."users" USING btree ("email");
CREATE UNIQUE INDEX "posts_slug_idx" ON "payload"."posts" USING btree ("slug");
CREATE INDEX "posts_hero_image_idx" ON "payload"."posts" USING btree ("hero_image_id");
CREATE INDEX "posts_updated_at_idx" ON "payload"."posts" USING btree ("updated_at");
CREATE INDEX "posts_created_at_idx" ON "payload"."posts" USING btree ("created_at");
CREATE INDEX "posts__status_idx" ON "payload"."posts" USING btree ("_status");
CREATE INDEX "_posts_v_parent_idx" ON "payload"."_posts_v" USING btree ("parent_id");
CREATE INDEX "_posts_v_version_version_slug_idx" ON "payload"."_posts_v" USING btree ("version_slug");
CREATE INDEX "_posts_v_version_version_hero_image_idx" ON "payload"."_posts_v" USING btree ("version_hero_image_id");
CREATE INDEX "_posts_v_version_version_updated_at_idx" ON "payload"."_posts_v" USING btree ("version_updated_at");
CREATE INDEX "_posts_v_version_version_created_at_idx" ON "payload"."_posts_v" USING btree ("version_created_at");
CREATE INDEX "_posts_v_version_version__status_idx" ON "payload"."_posts_v" USING btree ("version__status");
CREATE INDEX "_posts_v_created_at_idx" ON "payload"."_posts_v" USING btree ("created_at");
CREATE INDEX "_posts_v_updated_at_idx" ON "payload"."_posts_v" USING btree ("updated_at");
CREATE INDEX "_posts_v_latest_idx" ON "payload"."_posts_v" USING btree ("latest");
CREATE INDEX "media_updated_at_idx" ON "payload"."media" USING btree ("updated_at");
CREATE INDEX "media_created_at_idx" ON "payload"."media" USING btree ("created_at");
CREATE UNIQUE INDEX "media_filename_idx" ON "payload"."media" USING btree ("filename");
CREATE INDEX "media_sizes_thumbnail_sizes_thumbnail_filename_idx" ON "payload"."media" USING btree ("sizes_thumbnail_filename");
CREATE INDEX "media_sizes_hero_sizes_hero_filename_idx" ON "payload"."media" USING btree ("sizes_hero_filename");
CREATE UNIQUE INDEX "payload_kv_key_idx" ON "payload"."payload_kv" USING btree ("key");
CREATE INDEX "payload_locked_documents_global_slug_idx" ON "payload"."payload_locked_documents" USING btree ("global_slug");
CREATE INDEX "payload_locked_documents_updated_at_idx" ON "payload"."payload_locked_documents" USING btree ("updated_at");
CREATE INDEX "payload_locked_documents_created_at_idx" ON "payload"."payload_locked_documents" USING btree ("created_at");
CREATE INDEX "payload_locked_documents_rels_order_idx" ON "payload"."payload_locked_documents_rels" USING btree ("order");
CREATE INDEX "payload_locked_documents_rels_parent_idx" ON "payload"."payload_locked_documents_rels" USING btree ("parent_id");
CREATE INDEX "payload_locked_documents_rels_path_idx" ON "payload"."payload_locked_documents_rels" USING btree ("path");
CREATE INDEX "payload_locked_documents_rels_users_id_idx" ON "payload"."payload_locked_documents_rels" USING btree ("users_id");
CREATE INDEX "payload_locked_documents_rels_posts_id_idx" ON "payload"."payload_locked_documents_rels" USING btree ("posts_id");
CREATE INDEX "payload_locked_documents_rels_media_id_idx" ON "payload"."payload_locked_documents_rels" USING btree ("media_id");
CREATE INDEX "payload_preferences_key_idx" ON "payload"."payload_preferences" USING btree ("key");
CREATE INDEX "payload_preferences_updated_at_idx" ON "payload"."payload_preferences" USING btree ("updated_at");
CREATE INDEX "payload_preferences_created_at_idx" ON "payload"."payload_preferences" USING btree ("created_at");
CREATE INDEX "payload_preferences_rels_order_idx" ON "payload"."payload_preferences_rels" USING btree ("order");
CREATE INDEX "payload_preferences_rels_parent_idx" ON "payload"."payload_preferences_rels" USING btree ("parent_id");
CREATE INDEX "payload_preferences_rels_path_idx" ON "payload"."payload_preferences_rels" USING btree ("path");
CREATE INDEX "payload_preferences_rels_users_id_idx" ON "payload"."payload_preferences_rels" USING btree ("users_id");
CREATE INDEX "payload_migrations_updated_at_idx" ON "payload"."payload_migrations" USING btree ("updated_at");
CREATE INDEX "payload_migrations_created_at_idx" ON "payload"."payload_migrations" USING btree ("created_at");`)
}
export async function down({ db, payload, req }: MigrateDownArgs): Promise<void> {
await db.execute(sql`
DROP TABLE "payload"."users_sessions" CASCADE;
DROP TABLE "payload"."users" CASCADE;
DROP TABLE "payload"."posts" CASCADE;
DROP TABLE "payload"."_posts_v" CASCADE;
DROP TABLE "payload"."media" CASCADE;
DROP TABLE "payload"."payload_kv" CASCADE;
DROP TABLE "payload"."payload_locked_documents" CASCADE;
DROP TABLE "payload"."payload_locked_documents_rels" CASCADE;
DROP TABLE "payload"."payload_preferences" CASCADE;
DROP TABLE "payload"."payload_preferences_rels" CASCADE;
DROP TABLE "payload"."payload_migrations" CASCADE;
DROP TYPE "payload"."enum_posts_status";
DROP TYPE "payload"."enum__posts_v_version_status";`)
}
+9
View File
@@ -0,0 +1,9 @@
import * as migration_20260902_172826_initial from './20260902_172826_initial';
export const migrations = [
{
up: migration_20260902_172826_initial.up,
down: migration_20260902_172826_initial.down,
name: '20260902_172826_initial'
},
];
@@ -1,3 +1,5 @@
import { withPayload } from '@payloadcms/next/withPayload';
/** @type {import('next').NextConfig} */
const nextConfig = {
// Enable standalone output for Docker
@@ -86,6 +88,23 @@ const nextConfig = {
},
],
},
{
/*
* The admin panel and the CMS API must never be indexed.
*
* X-Robots-Tag, not just the robots.txt Disallow, for the same reason
* the staging rule above uses one: a Disallow blocks crawling, which
* is not indexing. A disallowed URL found from an external link can
* still be indexed without ever being fetched — and worse, blocking
* the crawl means the noindex is never seen.
*/
source: '/admin/:path*',
headers: [{ key: 'X-Robots-Tag', value: 'noindex, nofollow' }],
},
{
source: '/cms-api/:path*',
headers: [{ key: 'X-Robots-Tag', value: 'noindex, nofollow' }],
},
{
source: '/:path*',
headers: [
@@ -127,4 +146,4 @@ const nextConfig = {
},
};
module.exports = nextConfig;
export default withPayload(nextConfig);
+5184 -223
View File
File diff suppressed because it is too large. Load diff
+9 -2
View File
@@ -2,30 +2,38 @@
"name": "nextjs-app",
"version": "0.1.0",
"private": true,
"type": "module",
"description": "SchoolCompare Next.js Application",
"scripts": {
"dev": "next dev",
"build": "next build",
"start": "next start",
"typecheck": "tsc --noEmit",
"generate:importmap": "payload generate:importmap",
"test": "jest",
"test:watch": "jest --watch",
"test:coverage": "jest --coverage"
},
"dependencies": {
"@floating-ui/react": "^0.27.20",
"@payloadcms/db-postgres": "^3.88.0",
"@payloadcms/next": "^3.88.0",
"@payloadcms/richtext-lexical": "^3.88.0",
"@types/node": "^25.2.0",
"@types/react": "^19.2.10",
"@types/react-dom": "^19.2.3",
"chart.js": "^4.5.1",
"eslint": "^9.39.2",
"eslint-config-next": "^16.1.6",
"graphql": "^16.14.2",
"leaflet": "^1.9.4",
"next": "^16.1.6",
"payload": "^3.88.0",
"react": "^19.2.4",
"react-chartjs-2": "^5.3.1",
"react-dom": "^19.2.4",
"react-leaflet": "^5.0.0",
"sharp": "^0.35.4",
"typescript": "^5.9.3",
"zod": "^4.3.6"
},
@@ -36,7 +44,6 @@
"@types/jest": "^30.0.0",
"@types/leaflet": "^1.9.21",
"jest": "^30.2.0",
"jest-environment-jsdom": "^30.2.0",
"sharp": "^0.34.5"
"jest-environment-jsdom": "^30.2.0"
}
}
+443
View File
@@ -0,0 +1,443 @@
/* tslint:disable */
/* eslint-disable */
/**
* This file was automatically generated by Payload.
* DO NOT MODIFY IT BY HAND. Instead, modify your source Payload config,
* and re-run `payload generate:types` to regenerate this file.
*/
/**
* Supported timezones in IANA format.
*
* This interface was referenced by `Config`'s JSON-Schema
* via the `definition` "supportedTimezones".
*/
export type SupportedTimezones =
| 'Pacific/Midway'
| 'Pacific/Niue'
| 'Pacific/Honolulu'
| 'Pacific/Rarotonga'
| 'America/Anchorage'
| 'Pacific/Gambier'
| 'America/Los_Angeles'
| 'America/Tijuana'
| 'America/Denver'
| 'America/Phoenix'
| 'America/Chicago'
| 'America/Guatemala'
| 'America/New_York'
| 'America/Bogota'
| 'America/Caracas'
| 'America/Santiago'
| 'America/Buenos_Aires'
| 'America/Sao_Paulo'
| 'Atlantic/South_Georgia'
| 'Atlantic/Azores'
| 'Atlantic/Cape_Verde'
| 'Europe/London'
| 'Europe/Berlin'
| 'Africa/Lagos'
| 'Europe/Athens'
| 'Africa/Cairo'
| 'Europe/Moscow'
| 'Asia/Riyadh'
| 'Asia/Dubai'
| 'Asia/Baku'
| 'Asia/Karachi'
| 'Asia/Tashkent'
| 'Asia/Calcutta'
| 'Asia/Dhaka'
| 'Asia/Almaty'
| 'Asia/Jakarta'
| 'Asia/Bangkok'
| 'Asia/Shanghai'
| 'Asia/Singapore'
| 'Asia/Tokyo'
| 'Asia/Seoul'
| 'Australia/Brisbane'
| 'Australia/Sydney'
| 'Pacific/Guam'
| 'Pacific/Noumea'
| 'Pacific/Auckland'
| 'Pacific/Fiji';
export interface Config {
auth: {
users: UserAuthOperations;
};
blocks: {};
collections: {
users: User;
posts: Post;
media: Media;
'payload-kv': PayloadKv;
'payload-locked-documents': PayloadLockedDocument;
'payload-preferences': PayloadPreference;
'payload-migrations': PayloadMigration;
};
collectionsJoins: {};
collectionsSelect: {
users: UsersSelect<false> | UsersSelect<true>;
posts: PostsSelect<false> | PostsSelect<true>;
media: MediaSelect<false> | MediaSelect<true>;
'payload-kv': PayloadKvSelect<false> | PayloadKvSelect<true>;
'payload-locked-documents': PayloadLockedDocumentsSelect<false> | PayloadLockedDocumentsSelect<true>;
'payload-preferences': PayloadPreferencesSelect<false> | PayloadPreferencesSelect<true>;
'payload-migrations': PayloadMigrationsSelect<false> | PayloadMigrationsSelect<true>;
};
db: {
defaultIDType: number;
};
fallbackLocale: null;
globals: {};
globalsSelect: {};
locale: null;
widgets: {
collections: CollectionsWidget;
};
user: User;
jobs: {
tasks: unknown;
workflows: unknown;
};
}
export interface UserAuthOperations {
forgotPassword: {
email: string;
password: string;
};
login: {
email: string;
password: string;
};
registerFirstUser: {
email: string;
password: string;
};
unlock: {
email: string;
password: string;
};
}
/**
* This interface was referenced by `Config`'s JSON-Schema
* via the `definition` "users".
*/
export interface User {
id: number;
displayName: string;
updatedAt: string;
createdAt: string;
email: string;
resetPasswordToken?: string | null;
resetPasswordExpiration?: string | null;
salt?: string | null;
hash?: string | null;
loginAttempts?: number | null;
lockUntil?: string | null;
sessions?:
| {
id: string;
createdAt?: string | null;
expiresAt: string;
}[]
| null;
password?: string | null;
collection: 'users';
}
/**
* This interface was referenced by `Config`'s JSON-Schema
* via the `definition` "posts".
*/
export interface Post {
id: number;
title: string;
/**
* The URL segment. Never change it after publishing.
*/
slug: string;
publishedAt: string;
/**
* Shown on the index and used as the meta description.
*/
excerpt: string;
heroImage?: (number | null) | Media;
content: {
root: {
type: string;
children: {
type: any;
version: number;
[k: string]: unknown;
}[];
direction: ('ltr' | 'rtl') | null;
format: 'left' | 'start' | 'center' | 'right' | 'end' | 'justify' | '';
indent: number;
version: number;
};
[k: string]: unknown;
};
updatedAt: string;
createdAt: string;
_status?: ('draft' | 'published') | null;
}
/**
* This interface was referenced by `Config`'s JSON-Schema
* via the `definition` "media".
*/
export interface Media {
id: number;
/**
* Describe the image for screen readers.
*/
alt: string;
updatedAt: string;
createdAt: string;
url?: string | null;
thumbnailURL?: string | null;
filename?: string | null;
mimeType?: string | null;
filesize?: number | null;
width?: number | null;
height?: number | null;
focalX?: number | null;
focalY?: number | null;
sizes?: {
thumbnail?: {
url?: string | null;
width?: number | null;
height?: number | null;
mimeType?: string | null;
filesize?: number | null;
filename?: string | null;
};
hero?: {
url?: string | null;
width?: number | null;
height?: number | null;
mimeType?: string | null;
filesize?: number | null;
filename?: string | null;
};
};
}
/**
* This interface was referenced by `Config`'s JSON-Schema
* via the `definition` "payload-kv".
*/
export interface PayloadKv {
id: number;
key: string;
data:
| {
[k: string]: unknown;
}
| unknown[]
| string
| number
| boolean
| null;
}
/**
* This interface was referenced by `Config`'s JSON-Schema
* via the `definition` "payload-locked-documents".
*/
export interface PayloadLockedDocument {
id: number;
document?:
| ({
relationTo: 'users';
value: number | User;
} | null)
| ({
relationTo: 'posts';
value: number | Post;
} | null)
| ({
relationTo: 'media';
value: number | Media;
} | null);
globalSlug?: string | null;
user: {
relationTo: 'users';
value: number | User;
};
updatedAt: string;
createdAt: string;
}
/**
* This interface was referenced by `Config`'s JSON-Schema
* via the `definition` "payload-preferences".
*/
export interface PayloadPreference {
id: number;
user: {
relationTo: 'users';
value: number | User;
};
key?: string | null;
value?:
| {
[k: string]: unknown;
}
| unknown[]
| string
| number
| boolean
| null;
updatedAt: string;
createdAt: string;
}
/**
* This interface was referenced by `Config`'s JSON-Schema
* via the `definition` "payload-migrations".
*/
export interface PayloadMigration {
id: number;
name?: string | null;
batch?: number | null;
updatedAt: string;
createdAt: string;
}
/**
* This interface was referenced by `Config`'s JSON-Schema
* via the `definition` "users_select".
*/
export interface UsersSelect<T extends boolean = true> {
displayName?: T;
updatedAt?: T;
createdAt?: T;
email?: T;
resetPasswordToken?: T;
resetPasswordExpiration?: T;
salt?: T;
hash?: T;
loginAttempts?: T;
lockUntil?: T;
sessions?:
| T
| {
id?: T;
createdAt?: T;
expiresAt?: T;
};
}
/**
* This interface was referenced by `Config`'s JSON-Schema
* via the `definition` "posts_select".
*/
export interface PostsSelect<T extends boolean = true> {
title?: T;
slug?: T;
publishedAt?: T;
excerpt?: T;
heroImage?: T;
content?: T;
updatedAt?: T;
createdAt?: T;
_status?: T;
}
/**
* This interface was referenced by `Config`'s JSON-Schema
* via the `definition` "media_select".
*/
export interface MediaSelect<T extends boolean = true> {
alt?: T;
updatedAt?: T;
createdAt?: T;
url?: T;
thumbnailURL?: T;
filename?: T;
mimeType?: T;
filesize?: T;
width?: T;
height?: T;
focalX?: T;
focalY?: T;
sizes?:
| T
| {
thumbnail?:
| T
| {
url?: T;
width?: T;
height?: T;
mimeType?: T;
filesize?: T;
filename?: T;
};
hero?:
| T
| {
url?: T;
width?: T;
height?: T;
mimeType?: T;
filesize?: T;
filename?: T;
};
};
}
/**
* This interface was referenced by `Config`'s JSON-Schema
* via the `definition` "payload-kv_select".
*/
export interface PayloadKvSelect<T extends boolean = true> {
key?: T;
data?: T;
}
/**
* This interface was referenced by `Config`'s JSON-Schema
* via the `definition` "payload-locked-documents_select".
*/
export interface PayloadLockedDocumentsSelect<T extends boolean = true> {
document?: T;
globalSlug?: T;
user?: T;
updatedAt?: T;
createdAt?: T;
}
/**
* This interface was referenced by `Config`'s JSON-Schema
* via the `definition` "payload-preferences_select".
*/
export interface PayloadPreferencesSelect<T extends boolean = true> {
user?: T;
key?: T;
value?: T;
updatedAt?: T;
createdAt?: T;
}
/**
* This interface was referenced by `Config`'s JSON-Schema
* via the `definition` "payload-migrations_select".
*/
export interface PayloadMigrationsSelect<T extends boolean = true> {
name?: T;
batch?: T;
updatedAt?: T;
createdAt?: T;
}
/**
* This interface was referenced by `Config`'s JSON-Schema
* via the `definition` "collections_widget".
*/
export interface CollectionsWidget {
data?: {
[k: string]: unknown;
};
width: 'full';
}
/**
* This interface was referenced by `Config`'s JSON-Schema
* via the `definition` "auth".
*/
export interface Auth {
[k: string]: unknown;
}
declare module 'payload' {
export interface GeneratedTypes extends Config {}
}
+39
View File
@@ -0,0 +1,39 @@
import path from 'path';
import { fileURLToPath } from 'url';
import { buildConfig } from 'payload';
import { postgresAdapter } from '@payloadcms/db-postgres';
import { lexicalEditor } from '@payloadcms/richtext-lexical';
import sharp from 'sharp';
import { Users } from '@/collections/Users';
import { Posts } from '@/collections/Posts';
import { Media } from '@/collections/Media';
import { PAYLOAD_API_ROUTE, PAYLOAD_ADMIN_ROUTE } from '@/lib/payloadRoutes';
import { migrations } from '@/migrations';
const filename = fileURLToPath(import.meta.url);
const dirname = path.dirname(filename);
export default buildConfig({
admin: { user: Users.slug },
// Defined in lib/payloadRoutes.ts, which carries the reasoning and is what
// the test asserts. Never inline these — /api belongs to the FastAPI proxy.
routes: { api: PAYLOAD_API_ROUTE, admin: PAYLOAD_ADMIN_ROUTE },
collections: [Users, Posts, Media],
editor: lexicalEditor(),
secret: process.env.PAYLOAD_SECRET || '',
typescript: { outputFile: path.resolve(dirname, 'payload-types.ts') },
db: postgresAdapter({
pool: { connectionString: process.env.DATABASE_URL },
// Its own schema separates blog content from pipeline-managed school
// tables and Airflow metadata. The schema is created by the initial
// migration: schemaName says where tables go, it does not create anything.
//
// prodMigrations runs pending migrations during server init. Without it a
// production container connects to an empty schema and fails its first
// query with 42P01 — the adapter cannot self-create tables, because
// db-postgres/connect.js gates push on NODE_ENV !== 'production'.
schemaName: 'payload',
prodMigrations: migrations,
}),
sharp,
});
Loaded 100 of 106 files, more files were not shown because too many files have changed in this diff. Show more