Files
school_compare/docs/DEPLOY.md
TudorandClaude Opus 5 264edd2e3a
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m7s
PR Checks / Backend Smoke (pull_request) Successful in 10s
PR Checks / Build Backend (no push) (pull_request) Successful in 12s
PR Checks / Build Frontend (no push) (pull_request) Successful in 50s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 51s
PR Checks / AI Code Review (Claude) (pull_request) Failing after 2m58s
fix(airflow): a login that survives a container restart
The simple auth manager generates a random password on first start and
writes it to a file, so every restart of the api-server invalidated the
last one and the password had to be dug out of the container logs again.

The stack now writes that file itself from AIRFLOW_ADMIN_PASSWORD before
exec'ing the api-server. Airflow generates nothing when the file already
exists, so the login is whatever the stack environment says it is.

Written with python rather than echo, so json.dumps escapes a password
containing quotes, backslashes or non-ASCII correctly — verified against
`p@ss "wo\rd' £5`, which round-trips intact.

An unset AIRFLOW_ADMIN_PASSWORD raises KeyError and the container exits.
Falling back to a generated password would silently undo the point of the
change, and a compose-level `:?` gives the same refusal a readable reason.
This does mean the variable MUST be set in Portainer before the next
deploy of either stack.

Not affected by the two Docker gotchas in the upstream docs: this image
has no USER directive so it runs as root, and the file is rewritten from
the environment on every start rather than persisted on a volume.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvdDKvFFSZuMVDH5fEyTob
2026-08-31 17:38:42 +01:00

12 KiB

SDLC & Deployment Pipeline

SchoolCompare uses a two-stage deploy model on Gitea Actions with two human approvals. AI writes the code on feature branches; the first approval merges the PR, which deploys to staging and runs the E2E gate; the second approval — after manual testing on staging — promotes the exact same images to production via a manual workflow.

The flow

feature branch (AI-authored)
   │  PR to main                                    ← approval #1
   ▼
PR checks (.gitea/workflows/pr-checks.yml)
   typecheck + unit tests + backend smoke + image builds (no push)
   + Claude code review posted as a PR comment (severe findings fail the check)
   │  merge (branch protection requires green checks)
   ▼
Stage pipeline (.gitea/workflows/deploy.yml) — automatic
   1. build & push images        → tags sha-<sha>, staging
   2. staging Portainer webhook  → wait for staging health
   3. Playwright E2E journeys against staging      ← gate before human testing
   ▼
Manual testing on staging (stx.schoolcompare.co.uk)
   │  Actions → "Promote to Production (manual)"   ← approval #2
   ▼
Promote pipeline (.gitea/workflows/promote.yml) — manual dispatch
   1. resolve target sha (input, or latest main if empty)
   2. REFUSE unless that commit's "E2E Journeys against Staging" status is green
   3. retag sha-<sha> → :prod    (same bytes — build once, promote the image)
      previous :prod saved as :prod-previous
   4. prod Portainer webhook     → wait for prod health

Key principle: build once, promote the exact image. Production pins :prod, which only moves when a human runs the promote workflow — and the workflow only accepts commits that passed the staging E2E gate. Nothing tags :latest anymore.

Branch & PR workflow

  • main is protected: no direct pushes, PRs require green status checks.
  • All work (human or AI) happens on feature branches → PR to main.
  • Merging to main releases to staging only. Production moves only on the second approval. If staging or the E2E gate fails, fix forward — production is untouched either way.

Promotion granularity

Staging always runs the latest main. Promoting approves a state of main, not a single PR — if two PRs merged since the last promotion, they ship together. Test staging accordingly. To promote an older state, pass its commit SHA to the promote workflow (its images must still exist in the registry).

Staging quirk for manual testing: external /api is broken at the staging proxy — exercise API endpoints from the host, not via the public staging URL.

Environments

Production Staging
Portainer stack file docker-compose.portainer.yml docker-compose.portainer.staging.yml
Image tag :prod :staging
Container prefix sc_ / schoolcompare_ sc_staging_
Frontend macvlan IP 10.0.1.150 STAGING_FRONTEND_IP (default 10.0.1.151)
Postgres macvlan IP 10.0.1.189 STAGING_DB_IP (default 10.0.1.190)
Airflow UI port 8080 8081
Volumes stack-prefixed stack-prefixed (fully isolated)

Staging gets :staging images on every merge to main — even ones that later fail the E2E gate. That's the point: staging absorbs the risk.

Gitea repository secrets

Secret Purpose
REGISTRY_TOKEN push images to privaterepo.sitaru.org (already set)
CLAUDE_CODE_OAUTH_TOKEN Claude Code subscription auth for the PR review — generate with claude setup-token on your machine
PORTAINER_STAGING_WEBHOOK staging stack redeploy webhook URL
PORTAINER_PROD_WEBHOOK production stack redeploy webhook URL
STAGING_BASE_URL e.g. http://10.0.1.151:3000 — health poll + E2E target
PROD_BASE_URL e.g. http://10.0.1.150:3000 — post-promotion health poll

One-time setup checklist

  1. Create the staging stack in Portainer from docker-compose.portainer.staging.yml (stack name e.g. schoolcompare-staging). Set the same environment variables as prod plus STAGING_DB_IP / STAGING_FRONTEND_IP if the defaults clash.
  2. Enable webhooks on both stacks (Portainer → Stack → Webhook) and store the URLs as PORTAINER_STAGING_WEBHOOK / PORTAINER_PROD_WEBHOOK. Remove the old hardcoded webhook usage (now gone from the workflows).
  3. Add the remaining secrets listed above in Gitea → repo → Settings → Actions → Secrets.
  4. Protect main in Gitea → Settings → Branches: require PRs, require the pr-checks status checks (frontend, backend, builds, ai-review) to pass.
  5. Bootstrap staging data via Airflow (no prod dump — staging populates itself from source, exercising the pipeline image end-to-end):
    • Set AIRFLOW_ADMIN_PASSWORD in the stack environment first. The api-server refuses to start without it. Airflow's simple auth manager otherwise generates a password on first start and writes it to a file, so the login changes every time the container restarts; the stack writes that file itself from this variable instead. AIRFLOW_ADMIN_USER defaults to admin.
    • Open the staging Airflow UI (http://<host>:8081) and trigger, in order: school_data_daily, school_data_monthly_ofsted, then the manual-schedule school_data_annual_ees and school_data_annual_idaci.
    • First runs download from government sources (GIAS, Ofsted, EES, IDACI), run dbt, and sync Typesense — expect the initial backfill to take a while.
    • The scheduled DAGs then keep staging fresh exactly like prod.
  6. Switch the prod stack to :prod tags — the repo's docker-compose.portainer.yml is already updated; redeploy the prod stack from it. Until the first pipeline run promotes an image, tag the current images manually: docker buildx imagetools create -t <image>:prod <image>:latest for each of the three images.

Rollback

Re-run "Promote to Production (manual)" with the SHA of the last good commit (fastest, fully gated), or manually re-point the tags — every promotion first saves the outgoing :prod as :prod-previous:

for img in backend frontend pipeline; do
  docker buildx imagetools create \
    -t privaterepo.sitaru.org/tudor/school_compare-$img:prod \
    privaterepo.sitaru.org/tudor/school_compare-$img:prod-previous
done
curl -fsSk -X POST "$PORTAINER_PROD_WEBHOOK"

Or promote any older build directly: imagetools create -t <image>:prod <image>:sha-<shortsha>.

E2E suite

Lives in e2e/ (own package — CI installs it without the app's node_modules). Journeys: home + name search, postcode search, school detail, two-school comparison, rankings table. Run locally against any environment:

cd e2e && npm ci
BASE_URL=http://10.0.1.151:3000 npx playwright test

Tests assert data invariants (results exist, charts render), not exact numbers, so scheduled data refreshes don't break the gate.

AI code review

scripts/ci/ai_review.py pipes the PR diff through headless Claude Code (claude -p, authenticated with the subscription OAuth token — no API billing), posts the structured findings as a PR comment using the per-run token Gitea Actions provides automatically (secrets.GITEA_TOKEN — no setup needed), and fails the check only when a finding is rated severe (would break prod, leak data, or corrupt data). Minor findings are informational and never block a merge.

Rate limiting, and the Cloudflare gap

Two independent limits protect the API:

  • Per client, via slowapi, keyed on CF-Connecting-IP (falling back to X-Forwarded-For, then the peer address). 60/minute by default; /api/suggest gets 120/minute because typing is bursty.
  • Globally, via GlobalRateLimitMiddleware: a fixed 60-second window over all /api/ traffic, GLOBAL_RATE_LIMIT_PER_MINUTE (default 3000), independent of any client identity. Requests from 127.0.0.1 are exempt so the container healthcheck cannot be starved into a restart loop.

Open: the origin must only accept Cloudflare

CF-Connecting-IP is only meaningful for requests that actually reached the origin through Cloudflare, and the application cannot verify that they did. Anything able to reach the origin directly can set that header freely and, by rotating it, mint a fresh rate-limit bucket per request — defeating per-client limits on every endpoint.

The global ceiling bounds the damage to total origin capacity. It does not fix the underlying gap, and nothing in the code can. Closing it needs one of:

  • Authenticated Origin Pulls — Cloudflare presents a client certificate the origin requires, so non-Cloudflare traffic is refused at TLS.
  • An origin firewall restricted to Cloudflare's published IP ranges.

Until one is in place, treat per-client limits as protection against accidents and ordinary load, not against a determined caller.

Feature flags (Unleash)

Flag state lives in a self-hosted Unleash instance, deployed as its own Portainer stack from docker-compose.portainer.unleash.yml. It is separate from the application stacks on purpose — redeploying staging must not be able to disturb production's flags.

The flags themselves are declared in backend/flags.py. Unleash holds the state; the registry holds the list. A flag in the UI that is not in the registry is orphaned and nothing reads it.

First-time setup

  1. Deploy the stack in Portainer. Set UNLEASH_DB_PASSWORD, UNLEASH_ADMIN_PASSWORD and (optionally) UNLEASH_IP.

  2. Log in to the UI at http://<UNLEASH_IP>:4242 as admin.

  3. Create one client API token per environment:

    • schoolcompare-staging, environment development
    • schoolcompare-prod, environment production

    Client tokens, not admin tokens — the backend only reads.

  4. Put each token in the matching Portainer stack's UNLEASH_API_TOKEN variable, and set UNLEASH_URL to http://<UNLEASH_IP>:4242/api.

  5. Redeploy the application stacks.

Adding a flag to Unleash

Unleash does not create flags by itself. The SDK reads definitions from the server and never registers anything, and metrics for a flag the server has never heard of are discarded. So a flag declared in backend/flags.py will be evaluated on every request, stay False forever, and never appear in the UI until someone creates it there by hand.

For each flag in the registry, create one in Unleash with:

  • Name — character for character what backend/flags.py declares. snake_case, no hyphens or spaces. A typo produces a flag that looks correct in the UI and is read by nothing.
  • Type — Release. No strategies, constraints or variants: these are plain on/off switches, by design.

Turning a feature on

Toggle the flag in the environment matching the stack you mean: development for staging, production for prod. The token in each stack is scoped to one environment, so toggling the other one has no visible effect.

The SDK refreshes every 15 seconds, so the API reflects the change almost at once; the pages follow on their own schedule, below.

A flip reaches school pages within about five minutes and place pages within the hour. Next's ISR does the propagating — it revalidates a route at the lowest revalidate among that route's fetches, which is 300s for /school/[slug] and 3600s for the place pages. There is no webhook, and adding one would only be worth it if flips ever needed to be instant.

When Unleash is unreachable

Every flag evaluates to False and the site serves as though nothing were switched on. That is deliberate — an unfinished feature staying hidden is the safe direction — but it means a released feature disappears if a backend container cold-starts with an empty cache while Unleash is down. The SDK's disk cache is on a named volume so restarts keep last-known state, and flags are removed from the code within 90 days (enforced by a test), which bounds how long any feature is exposed to this.

If UNLEASH_URL is unset, every flag is False and no connection is attempted. That is the correct behaviour for local development and CI, and it means the test suites need no flag server.