Staging health polling asked only whether something answered HTTP 200 at
the base URL. It could not tell the new deployment from the old one, so
journeys could pass against the previous release, and concurrent merges
could move the staging tags underneath a run in flight.
Each staging run now mints a build ID and stamps all three images with
the commit and that ID, as labels and — for frontend and backend — as a
build-time JSON file that environment overrides cannot rewrite.
/release.json reports both identities uncached, and scripts/ci/release.py
polls for the expected pair before and after the journeys. Only then are
the captured build digests tagged verified-<sha>.
Promotion resolves those verified tags to immutable digests, revalidates
their labels, and refuses a mixed or incomplete set before any :prod tag
moves. The whole staging workflow shares one concurrency group with
cancellation disabled, so releases serialise.
The scripts are stdlib-only and unit-tested against mocked registry and
HTTP behaviour; PR checks now run the pipeline and CI suites too. The
runbook records what this cannot prove locally, and that the first
rollout needs a commit built by this workflow.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The simple auth manager generates a random password on first start and
writes it to a file, so every restart of the api-server invalidated the
last one and the password had to be dug out of the container logs again.
The stack now writes that file itself from AIRFLOW_ADMIN_PASSWORD before
exec'ing the api-server. Airflow generates nothing when the file already
exists, so the login is whatever the stack environment says it is.
Written with python rather than echo, so json.dumps escapes a password
containing quotes, backslashes or non-ASCII correctly — verified against
`p@ss "wo\rd' £5`, which round-trips intact.
An unset AIRFLOW_ADMIN_PASSWORD raises KeyError and the container exits.
Falling back to a generated password would silently undo the point of the
change, and a compose-level `:?` gives the same refusal a readable reason.
This does mean the variable MUST be set in Portainer before the next
deploy of either stack.
Not affected by the two Docker gotchas in the upstream docs: this image
has no USER directive so it runs as root, and the file is rewritten from
the environment on every start rather than persisted on a volume.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvdDKvFFSZuMVDH5fEyTob
Code review, both findings valid.
The design doc claimed Cloudflare "replaces the header, so a browser
cannot forge it", and that only the X-Forwarded-For fallback was
forgeable. That is true only for traffic that actually passed through
Cloudflare, and nothing in this process can verify that it did. Reaching
the origin directly, both headers are equally attacker-controlled — and
rotating CF-Connecting-IP mints a fresh rate-limit bucket per request,
defeating per-client limits on every endpoint including the
DataFrame-heavy /api/schools. Against abuse that is worse than the
shared bucket it replaced, which at least capped everyone together.
So the ceiling comes back. I dropped it earlier arguing it belonged at
Cloudflare; that argument assumed the keying was sound, and it is not.
GlobalRateLimitMiddleware counts all /api/ traffic in a fixed window
against a total, independent of client identity, outermost so it refuses
before any work happens. Written by hand because slowapi cannot express
a global cap: default_limits and application_limits are both keyed by
key_func, and the latter needs middleware this app does not install.
It does not make the header trustworthy — it makes trusting it
survivable. The real fix is Authenticated Origin Pulls or an origin
firewall, now documented in DEPLOY.md as the open gap it is.
127.0.0.1 is exempt: the healthcheck curls localhost from inside the
container, and starving it would restart the container and turn a load
spike into an outage loop. Keyed on the peer address, never the Host
header, which the caller sets.
Second finding: suggest_schools_typesense promised "never raises" while
the parsing loop sat outside the try, so int(None) on a malformed
document would have made a keystroke a 500. The loop now skips bad rows
rather than dropping the whole list — and a hit with no document no
longer becomes a suggestion pointing at /school/0.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
The runbook said a flag 'appears in the Unleash UI after the backend has
evaluated it once'. That is wrong. SDKs read definitions from the server
and never register anything, and metrics for an unknown flag are
discarded — so a declared flag is evaluated on every request, stays
False forever, and never shows up until someone creates it by hand.
Found the way these things usually are: staging had been running the
flag code for a while and the UI was still empty.
Also names the environment trap while here — each stack's token is
scoped to one environment, so toggling the other does nothing visible
and looks like the flag is broken.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
Its own Portainer stack, belonging to neither application stack: a
staging redeploy must not be able to disturb production's flag state.
One instance serves both. OSS Unleash ships development and production
environments with environment-scoped client tokens, so the same flag
holds independent state in each — which is what lets a feature be on in
staging, where the E2E journeys exercise it, while production stays dark.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
Removes the 'What Parents Say' section and all supporting elements:
Frontend:
- Drop the OfstedParentView type, the parent_view field, the survey
section and the 'X% would recommend' callouts in the primary and
secondary detail views, the Parents nav item, and the parent-view CSS.
Backend:
- Remove the FactParentView model, its loading in data_loader, and
parent_view from the school-details API response.
- Bump SCHEMA_VERSION to 6 and add an idempotent drop step
(DROP TABLE IF EXISTS marts.fact_parent_view) to the CLI migration;
add scripts/sql/drop_fact_parent_view.sql to apply directly to the
dbt-owned marts DBs on staging and prod.
Pipeline:
- Delete the stg_parent_view + fact_parent_view dbt models and their
source/schema entries, the tap-uk-parent-view Meltano extractor, and
the monthly Parent View DAG; drop it from the Dockerfile and the
staging bootstrap docs.
The rest of dbt (which builds every mart the app reads) is untouched.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- ai_review.py now pipes the diff through headless Claude Code (claude -p,
--output-format json) authenticated with CLAUDE_CODE_OAUTH_TOKEN from
'claude setup-token' — subscription auth, no Anthropic API billing
- stdlib-only script (urllib instead of requests/anthropic)
- PR comments posted with the existing REGISTRY_TOKEN secret; the separate
GITEA_TOKEN secret is no longer needed
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PqGhF93UrpDNvXBLMjJENL
- pr-checks.yml: PR gate — frontend typecheck+jest, backend import smoke,
image builds (no push), Claude AI review posted as PR comment (severe
findings block merge)
- deploy.yml (replaces build-and-push.yml): merge to main builds+pushes
images tagged sha-<sha>/staging, deploys the staging Portainer stack via
webhook, runs Playwright E2E journeys against staging, then retags the
verified images :prod (previous kept as :prod-previous) and deploys prod
- docker-compose.portainer.staging.yml: second Portainer stack — :staging
images, sc_staging_* names, own macvlan IPs, Airflow on 8081; data
bootstrapped from source via the staging Airflow DAGs
- prod compose now pins :prod instead of :latest (only the promotion step
moves it; :latest is no longer published)
- e2e/: 6 Playwright journeys (search, postcode, detail, compare, rankings)
driven by BASE_URL — the promotion gate
- scripts/ci/ai_review.py: Claude review with structured JSON findings
- docs/DEPLOY.md: full SDLC doc incl. one-time setup checklist and rollback
- replaced removed 'next lint' with tsc typecheck; fixed stale jest tests
(slug URLs, N/A formatting, stable trend, fake-timer setup)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PqGhF93UrpDNvXBLMjJENL