Files
TudorandClaude Opus 5 cd6a45bf7d refactor: rename similar → nearby, so the code says what the section does
The section ranks on distance and is headed "Other schools nearby", but every
identifier still called it "similar" — the exact drift that leaves a later
reader trusting a name over the behaviour.

Mechanical: files, the module, the payload key, the type, the components, the
prop. No behaviour change; the suites are unchanged in count and still green.
Free to do now because #150 has not merged, so the payload key rename needs no
lockstep deploy. Uses of "similar" that are ordinary English — progress
measures compared to similar pupils, and unrelated comments — are untouched.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-22 13:06:32 +01:00

126 lines
7.0 KiB
Markdown

# Architecture
This describes the implementation as reviewed on 2026-09-14. It distinguishes
current behaviour from improvements still to be implemented.
## Request flow
```text
Browser → Next.js public routes
├─ /api/* proxy → FastAPI → cached DataFrames / PostgreSQL marts
│ ├─ Typesense (search and suggestions)
│ └─ postcodes.io (postcode lookup)
└─ /admin, /cms-api, /blog → Payload → payload schema + media volume
Next.js server rendering → FastAPI directly through FASTAPI_URL
```
`nextjs-app/lib/api.ts` contains typed fetch wrappers and revalidation defaults.
The proxy is `nextjs-app/app/(frontend)/api/[...path]/route.ts`. Payload uses
`/cms-api` so its routes do not collide with the FastAPI proxy. The proxy denies
`/api/flags`; server-side rendering reads flags directly from FastAPI.
## Data ownership
| Layer | Owner and role |
|---|---|
| Source data | GIAS, DfE EES, Ofsted, finance, deprivation and council admission-distance sources |
| `raw` | Singer taps and the PostgreSQL target configured in `pipeline/meltano.yml` |
| Staging/intermediate/marts | dbt models in `pipeline/transform`; marts are materialized tables |
| `marts.dim_school`, `marts.dim_location` | School identity and location, filtered to supported England establishments |
| `marts.fact_*` | Performance and supplementary datasets; coverage and years vary |
| Typesense `schools` alias | Search documents built by `pipeline/scripts/sync_typesense.py` |
| `payload` | CMS collections and migrations in `nextjs-app/`; independent of dbt |
| Media volume | Uploaded blog media; requires backup and cannot be regenerated from school datasets |
`backend/models.py` maps existing marts for reading. It does not create the school
schema. There is no startup schema-version migration or CSV reimport. Payload's
`nextjs-app/migrations/` is active and must not be confused with the removed
legacy backend migration code.
Coordinates normally come from GIAS British National Grid coordinates transformed
by PostGIS in `dim_location.sql`. `pipeline/scripts/geocode_postcodes.py` is a
manual fallback utility, not a task wired into the current school-data DAG.
Backend postcode searches also use postcodes.io; that lookup does not populate
school coordinates in the database.
## Backend boundaries
- `app.py`: routes, middleware, search filtering, sitemap/place publication and response assembly.
- `data_loader.py`: SQL loading, process-local DataFrame caches, Typesense calls,
postcode lookups, supplementary queries and benchmark calculation.
- `database.py`: synchronous SQLAlchemy engine and sessions.
- `schemas.py`: metric definitions, column mappings and display metadata; despite
its name this is not a collection of Pydantic API response models.
- `places.py` and `localities.py`: place registry and curated locality information.
- `flags.py`: Unleash-backed feature flags, disabled when no server is configured.
- `gias_codes.py` / `ofsted_codes.py`: source-code translation and display rules.
Search starts from a cached latest-row-per-school snapshot. Detail pages read
history from the full DataFrame and supplementary data from marts. Comparisons
batch supplementary queries across selected URNs. Async routes still contain
synchronous dependency calls; a fully asynchronous database layer is not present.
## Frontend boundaries
`app/(frontend)` owns the public root layout and pages. `app/(payload)` owns the
CMS root layout. Do not add a shared `app/layout.tsx`: these groups deliberately
have separate root layouts. Root metadata files remain in `app/`.
Server pages fetch initial data and pass it to client views. Client state uses
React hooks, URL search parameters and the comparison context/localStorage.
There is no SWR dependency. Leaflet maps are loaded through dynamic wrappers;
Chart.js renders performance and comparison charts.
`components/school/` contains detail sections, with section decisions and data
preparation in `lib/schoolSections.ts`. The nearby-schools section is selected in
`backend/nearby_schools.py` — hard filters decide eligibility (phase, provision,
selectivity, gender) and distance alone decides the order, capped per phase —
and served on `/api/schools/{urn}`. Its rules are presentation logic,
deliberately kept out of `marts.*` so they can be tuned by deploy rather than by
pipeline run. `lib/types.ts` contains manually maintained
API types. `payload-types.ts` and the Payload import map are generated artifacts.
## Publication and caching today
1. Airflow DAGs extract and validate source data, then run selected dbt builds.
2. Relevant DAGs rebuild Typesense and swap the `schools` alias.
3. They call `POST /api/admin/reload` with `X-API-Key`. It builds and validates
replacement DataFrames, places, reverse membership and sitemaps off the request
loop, then publishes them together. Failure returns 503 and preserves live data.
4. A separate weekly sitemap DAG can regenerate the derived publication from the
current DataFrame without clearing the live registry first.
GIAS is scheduled daily, Ofsted monthly, and annual datasets are manually
triggered. The DAG definitions are authoritative for selectors and dependencies.
Caches exist in several independent layers: backend DataFrames and registries,
backend HTTP Cache-Control/ETags, Next.js fetch/page revalidation, and browser or
shared HTTP caches where configured. Place fetches request a one-week revalidation
interval. HTTP ETags are computed after route execution, not before database work.
Typesense publication validates every import response and the final document
count before switching aliases. A session-scoped PostgreSQL advisory lock
serialises index reads/publication across DAGs. The previous collection remains
available for rollback; old unaliased collections are pruned after success.
Failed drafts are retained until a later successful cleanup, because an uncertain
alias-update response must never cause deletion of a potentially live index.
The backend snapshot swap is process-local and assumes the current single-worker
deployment. It is not an atomic transaction spanning PostgreSQL marts, Typesense
and Next.js caches. Next.js caches are not explicitly purged by the pipeline.
School search retrieves a relevance-ordered candidate prefix (currently capped at
1,000 URNs) before applying API filters. This keeps scoped searches useful while
putting a hard ceiling on Typesense round trips; only a dependency failure invokes
substring fallback, not a valid empty match set.
## Deployment references
See [DEPLOY.md](DEPLOY.md). PR checks include frontend typechecking/tests, backend
unit tests, image builds and AI review. Staging journeys run after merging.
Staging runs are serialised across builds, deployment and E2E. Build-stamped
frontend/backend identities are checked before and after journeys. Only then are
the captured image digests marked verified. Promotion resolves and validates the
complete verified image set before retagging production. See the runbook for
first-rollout requirements and remaining integration checks.