Files
school_compare/docs/DEPLOY.md
TudorandClaude Opus 5 64ae71d7ab
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m11s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 17s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m17s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 11s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 1m4s
fix(ci): make a failed release check say what it actually saw
The staging poller swallowed every failure identically, so a run that
timed out told us only that the expected release never appeared — not
whether the proxy refused us, the endpoint was down, or the containers
were still serving an older build. The public staging proxy also answers
403 to urllib's default user agent while the release endpoint is healthy,
which looked exactly like a deployment that never arrived.

Identify the poller, and report each distinct observation once: HTTP
status, connection failure type, invalid JSON, or the release identities
actually reported. The timeout error carries the last observation and the
identity it wanted. Responses and the base URL stay out of the logs —
only validated sha/build_id fields are echoed back.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 16:02:14 +01:00

308 lines
15 KiB
Markdown

# SDLC & Deployment Pipeline
SchoolCompare uses a two-stage deploy model on Gitea Actions with two human
approvals. AI writes the code on feature branches; the first approval merges
the PR, which deploys to staging and runs the E2E gate; the second approval —
after manual testing on staging — promotes the exact same images to
production via a manual workflow.
## The flow
```
feature branch (AI-authored)
│ PR to main ← approval #1
▼
PR checks (.gitea/workflows/pr-checks.yml)
typecheck + unit tests + backend smoke + image builds (no push)
+ Claude code review posted as a PR comment (severe findings fail the check)
│ merge (branch protection requires green checks)
▼
Stage pipeline (.gitea/workflows/deploy.yml) — automatic
1. build & push images → tags sha-<sha>, staging
2. staging Portainer webhook → verify frontend/backend SHA + build ID
3. Playwright E2E journeys against staging ← gate before human testing
4. verify identity again; tag tested digests verified-<full-sha>
▼
Manual testing on staging (stx.schoolcompare.co.uk)
│ Actions → "Promote to Production (manual)" ← approval #2
▼
Promote pipeline (.gitea/workflows/promote.yml) — manual dispatch
1. resolve target sha (input, or latest main if empty)
2. REFUSE unless that commit's "E2E Journeys against Staging" status is green
3. resolve verified-<full-sha> digests, validate labels, retag digests → :prod
previous :prod saved as :prod-previous
4. prod Portainer webhook → verify expected SHA + build ID
```
Key principle: **build once, promote the exact image**. Production pins `:prod`,
which only moves when a human runs the promote workflow — and the workflow
only accepts commits that passed the staging E2E gate and have a complete verified
image set. Nothing tags `:latest`
anymore.
## Branch & PR workflow
- `main` is protected: no direct pushes, PRs require green status checks.
- All work (human or AI) happens on feature branches → PR to `main`.
- Merging to `main` releases **to staging only**. Production moves only on
the second approval. If staging or the E2E gate fails, fix forward —
production is untouched either way.
## Promotion granularity
Staging always runs the latest `main`. Promoting approves a *state of main*,
not a single PR — if two PRs merged since the last promotion, they ship
together. Test staging accordingly. To promote an older state, pass its
commit SHA to the promote workflow (its images must still exist in the
registry).
Staging quirk for manual testing: external `/api` is broken at the staging
proxy — exercise API endpoints from the host, not via the public staging URL.
## Environments
| | Production | Staging |
|---|---|---|
| Portainer stack file | `docker-compose.portainer.yml` | `docker-compose.portainer.staging.yml` |
| Image tag | `:prod` | `:staging` |
| Container prefix | `sc_` / `schoolcompare_` | `sc_staging_` |
| Frontend macvlan IP | 10.0.1.150 | `STAGING_FRONTEND_IP` (default 10.0.1.151) |
| Postgres macvlan IP | 10.0.1.189 | `STAGING_DB_IP` (default 10.0.1.190) |
| Airflow UI port | 8080 | 8081 |
| Volumes | stack-prefixed | stack-prefixed (fully isolated) |
Staging gets `:staging` images on every merge to main — even ones that later
fail the E2E gate. That's the point: staging absorbs the risk.
## Gitea repository secrets
| Secret | Purpose |
|---|---|
| `REGISTRY_TOKEN` | push images to privaterepo.sitaru.org (already set) |
| `CLAUDE_CODE_OAUTH_TOKEN` | Claude Code subscription auth for the PR review — generate with `claude setup-token` on your machine |
| `PORTAINER_STAGING_WEBHOOK` | staging stack redeploy webhook URL |
| `PORTAINER_PROD_WEBHOOK` | production stack redeploy webhook URL |
| `STAGING_BASE_URL` | e.g. `http://10.0.1.151:3000` — health poll + E2E target |
| `PROD_BASE_URL` | e.g. `http://10.0.1.150:3000` — post-promotion health poll |
## One-time setup checklist
1. **Create the staging stack** in Portainer from
`docker-compose.portainer.staging.yml` (stack name e.g.
`schoolcompare-staging`). Set the same environment variables as prod plus
`STAGING_DB_IP` / `STAGING_FRONTEND_IP` if the defaults clash.
2. **Enable webhooks** on both stacks (Portainer → Stack → Webhook) and store
the URLs as `PORTAINER_STAGING_WEBHOOK` / `PORTAINER_PROD_WEBHOOK`. Remove
the old hardcoded webhook usage (now gone from the workflows).
3. **Add the remaining secrets** listed above in Gitea → repo → Settings →
Actions → Secrets.
4. **Protect `main`** in Gitea → Settings → Branches: require PRs, require the
pr-checks status checks (frontend, backend, builds, ai-review) to pass.
5. **Bootstrap staging data via Airflow** (no prod dump — staging populates
itself from source, exercising the pipeline image end-to-end):
- Set `AIRFLOW_ADMIN_PASSWORD` in the stack environment first. The
api-server refuses to start without it. Airflow's simple auth manager
otherwise generates a password on first start and writes it to a file, so
the login changes every time the container restarts; the stack writes that
file itself from this variable instead. `AIRFLOW_ADMIN_USER` defaults to
`admin`.
- Open the staging Airflow UI (`http://<host>:8081`) and trigger, in order:
`school_data_daily`, `school_data_monthly_ofsted`, then the manual-schedule
`school_data_annual_ees` and `school_data_annual_idaci`.
- First runs download from government sources (GIAS, Ofsted, EES, IDACI),
run dbt, and sync Typesense — expect the initial backfill to take a while.
- The scheduled DAGs then keep staging fresh exactly like prod.
6. **Switch the prod stack to `:prod` tags** — the repo's
`docker-compose.portainer.yml` is already updated; redeploy the prod stack
from it. Until the first pipeline run promotes an image, tag the current
images manually: `docker buildx imagetools create -t <image>:prod <image>:latest`
for each of the three images.
## Rollback
Re-run "Promote to Production (manual)" with the SHA of the last good commit
(fastest, fully gated), or manually re-point the tags — every promotion first
saves the outgoing `:prod` as `:prod-previous`:
```bash
for img in backend frontend pipeline; do
docker buildx imagetools create \
-t privaterepo.sitaru.org/tudor/school_compare-$img:prod \
privaterepo.sitaru.org/tudor/school_compare-$img:prod-previous
done
curl -fsSk -X POST "$PORTAINER_PROD_WEBHOOK"
```
Or promote any older build directly: `imagetools create -t <image>:prod <image>:sha-<shortsha>`.
## E2E suite
Lives in `e2e/` (own package — CI installs it without the app's node_modules).
Journeys: home + name search, postcode search, school detail, two-school
comparison, rankings table. Run locally against any environment:
```bash
cd e2e && npm ci
BASE_URL=http://10.0.1.151:3000 npx playwright test
```
Tests assert data invariants (results exist, charts render), not exact
numbers, so scheduled data refreshes don't break the gate.
## AI code review
`scripts/ci/ai_review.py` pipes the PR diff through headless Claude Code
(`claude -p`, authenticated with the subscription OAuth token — no API
billing), posts the structured findings as a PR comment using the per-run
token Gitea Actions provides automatically (`secrets.GITEA_TOKEN` — no setup
needed), and fails the check only when a finding is rated
**severe** (would break prod, leak data, or corrupt data). Minor findings are
informational and never block a merge.
## Rate limiting, and the Cloudflare gap
Two independent limits protect the API:
- **Per client**, via slowapi, keyed on `CF-Connecting-IP` (falling back to
`X-Forwarded-For`, then the peer address). 60/minute by default;
`/api/suggest` gets 120/minute because typing is bursty.
- **Globally**, via `GlobalRateLimitMiddleware`: a fixed 60-second window over
all `/api/` traffic, `GLOBAL_RATE_LIMIT_PER_MINUTE` (default 3000),
independent of any client identity. Requests from `127.0.0.1` are exempt so
the container healthcheck cannot be starved into a restart loop.
### Open: the origin must only accept Cloudflare
`CF-Connecting-IP` is only meaningful for requests that actually reached the
origin through Cloudflare, and **the application cannot verify that they did**.
Anything able to reach the origin directly can set that header freely and, by
rotating it, mint a fresh rate-limit bucket per request — defeating per-client
limits on every endpoint.
The global ceiling bounds the damage to total origin capacity. It does not fix
the underlying gap, and nothing in the code can. Closing it needs one of:
- **Authenticated Origin Pulls** — Cloudflare presents a client certificate the
origin requires, so non-Cloudflare traffic is refused at TLS.
- **An origin firewall** restricted to Cloudflare's published IP ranges.
Until one is in place, treat per-client limits as protection against accidents
and ordinary load, not against a determined caller.
## Feature flags (Unleash)
Flag state lives in a self-hosted Unleash instance, deployed as its own
Portainer stack from `docker-compose.portainer.unleash.yml`. It is separate
from the application stacks on purpose — redeploying staging must not be able
to disturb production's flags.
The flags themselves are declared in `backend/flags.py`. Unleash holds the
state; the registry holds the list. A flag in the UI that is not in the
registry is orphaned and nothing reads it.
### First-time setup
1. Deploy the stack in Portainer. Set `UNLEASH_DB_PASSWORD`,
`UNLEASH_ADMIN_PASSWORD` and (optionally) `UNLEASH_IP`.
2. Log in to the UI at `http://<UNLEASH_IP>:4242` as `admin`.
3. Create one **client** API token per environment:
- `schoolcompare-staging`, environment **development**
- `schoolcompare-prod`, environment **production**
Client tokens, not admin tokens — the backend only reads.
4. Put each token in the matching Portainer stack's `UNLEASH_API_TOKEN`
variable, and set `UNLEASH_URL` to `http://<UNLEASH_IP>:4242/api`.
5. Redeploy the application stacks.
### Adding a flag to Unleash
**Unleash does not create flags by itself.** The SDK reads definitions from the
server and never registers anything, and metrics for a flag the server has
never heard of are discarded. So a flag declared in `backend/flags.py` will be
evaluated on every request, stay `False` forever, and never appear in the UI
until someone creates it there by hand.
For each flag in the registry, create one in Unleash with:
- **Name** — character for character what `backend/flags.py` declares.
snake_case, no hyphens or spaces. A typo produces a flag that looks correct
in the UI and is read by nothing.
- **Type** — Release. No strategies, constraints or variants: these are plain
on/off switches, by design.
### Turning a feature on
Toggle the flag in the environment matching the stack you mean: **development**
for staging, **production** for prod. The token in each stack is scoped to one
environment, so toggling the other one has no visible effect.
The SDK refreshes every 15 seconds, so the API reflects the change almost at
once; the pages follow on their own schedule, below.
A flip reaches school pages within about five minutes and place pages within
the hour. Next's ISR does the propagating — it revalidates a route at the
*lowest* `revalidate` among that route's fetches, which is 300s for
`/school/[slug]` and 3600s for the place pages. There is no webhook, and
adding one would only be worth it if flips ever needed to be instant.
### When Unleash is unreachable
Every flag evaluates to `False` and the site serves as though nothing were
switched on. That is deliberate — an unfinished feature staying hidden is the
safe direction — but it means a *released* feature disappears if a backend
container cold-starts with an empty cache while Unleash is down. The SDK's
disk cache is on a named volume so restarts keep last-known state, and flags
are removed from the code within 90 days (enforced by a test), which bounds
how long any feature is exposed to this.
If `UNLEASH_URL` is unset, every flag is `False` and no connection is
attempted. That is the correct behaviour for local development and CI, and it
means the test suites need no flag server.
## Release identity and the P1 reliability gate
Every staging run creates a random build ID before building its three images.
Each image carries the commit and build ID as labels. Frontend/backend images
also contain a build-time JSON file; environment overrides cannot rewrite it.
`/release.json` returns both identities with `Cache-Control: no-store`. It fails
with 503 when either identity cannot be read. FastAPI's internal endpoint is
`/api/release`.
The entire staging workflow shares one concurrency group, with cancellation
disabled. This needs Gitea 1.26 or newer, where workflow concurrency is supported
([release notes](https://blog.gitea.com/release-of-1.26.0/)); the configured server
reported 1.27.3 during this change. Do not run the workflow on an older server
that ignores the concurrency key. Manual deployments outside this workflow must
also avoid changing staging during journeys.
The gate checks both identities before and after Playwright. It then validates
labels on the captured build output digests and tags them `verified-<full-sha>`.
The manual promotion script resolves all three verified tags to immutable digests
and confirms one matching commit/build ID before moving any `:prod` tag. It polls
production for that same identity using a locally saved release manifest.
A registry error can still interrupt the three tag writes; the Portainer webhook
only runs after successful promotion, and rerunning promotion resolves the full
verified set again. There is no cross-registry atomic tag transaction.
**First rollout:** old green commits without verified tags/build identities are
not promotable through this gate. Build and test a commit containing the new
workflow first. The release route must be reachable through the configured
`STAGING_BASE_URL`/`PROD_BASE_URL`; it deliberately avoids the public staging
`/api` proxy limitation. No new deployment secret is required.
`scripts/ci/release.py` implements identity polling and digest verification.
The poller identifies itself as `SchoolCompare-Release-Check/1.0`: the public
staging proxy has returned HTTP 403 to Python's default urllib user agent even
while the release endpoint was healthy. It logs changes in HTTP/connection
failures or observed release identities, and includes the last observation in
the timeout error. If verification fails, use that observation to distinguish
proxy rejection (403), an unavailable release endpoint (503), and containers
still reporting an older SHA/build ID. Check the configured base URL from the
CI runner; a successful request from another machine does not establish runner
connectivity. Do not bypass identity verification to unblock a deployment.
Its mocked tests run in PR checks alongside backend and index-publication tests.
The new Playwright journeys also check deployed identity and stale pagination.
Local unit checks do not validate registry credentials, Portainer behaviour,
proxy routing or a deployed image; those require the staging run. Production
promotion remains a separate human action.