2026-07-03 06:50:40 +01:00
|
|
|
# SDLC & Deployment Pipeline
|
|
|
|
|
|
2026-07-13 08:38:37 +01:00
|
|
|
SchoolCompare uses a two-stage deploy model on Gitea Actions with two human
|
|
|
|
|
approvals. AI writes the code on feature branches; the first approval merges
|
|
|
|
|
the PR, which deploys to staging and runs the E2E gate; the second approval —
|
|
|
|
|
after manual testing on staging — promotes the exact same images to
|
|
|
|
|
production via a manual workflow.
|
2026-07-03 06:50:40 +01:00
|
|
|
|
|
|
|
|
## The flow
|
|
|
|
|
|
|
|
|
|
```
|
|
|
|
|
feature branch (AI-authored)
|
2026-07-13 08:38:37 +01:00
|
|
|
│ PR to main ← approval #1
|
2026-07-03 06:50:40 +01:00
|
|
|
▼
|
|
|
|
|
PR checks (.gitea/workflows/pr-checks.yml)
|
|
|
|
|
typecheck + unit tests + backend smoke + image builds (no push)
|
|
|
|
|
+ Claude code review posted as a PR comment (severe findings fail the check)
|
|
|
|
|
│ merge (branch protection requires green checks)
|
|
|
|
|
▼
|
2026-07-13 08:38:37 +01:00
|
|
|
Stage pipeline (.gitea/workflows/deploy.yml) — automatic
|
2026-07-03 06:50:40 +01:00
|
|
|
1. build & push images → tags sha-<sha>, staging
|
|
|
|
|
2. staging Portainer webhook → wait for staging health
|
2026-07-13 08:38:37 +01:00
|
|
|
3. Playwright E2E journeys against staging ← gate before human testing
|
|
|
|
|
▼
|
|
|
|
|
Manual testing on staging (stx.schoolcompare.co.uk)
|
|
|
|
|
│ Actions → "Promote to Production (manual)" ← approval #2
|
|
|
|
|
▼
|
|
|
|
|
Promote pipeline (.gitea/workflows/promote.yml) — manual dispatch
|
|
|
|
|
1. resolve target sha (input, or latest main if empty)
|
|
|
|
|
2. REFUSE unless that commit's "E2E Journeys against Staging" status is green
|
|
|
|
|
3. retag sha-<sha> → :prod (same bytes — build once, promote the image)
|
2026-07-03 06:50:40 +01:00
|
|
|
previous :prod saved as :prod-previous
|
2026-07-13 08:38:37 +01:00
|
|
|
4. prod Portainer webhook → wait for prod health
|
2026-07-03 06:50:40 +01:00
|
|
|
```
|
|
|
|
|
|
|
|
|
|
Key principle: **build once, promote the exact image**. Production pins `:prod`,
|
2026-07-13 08:38:37 +01:00
|
|
|
which only moves when a human runs the promote workflow — and the workflow
|
|
|
|
|
only accepts commits that passed the staging E2E gate. Nothing tags `:latest`
|
2026-07-03 06:50:40 +01:00
|
|
|
anymore.
|
|
|
|
|
|
|
|
|
|
## Branch & PR workflow
|
|
|
|
|
|
|
|
|
|
- `main` is protected: no direct pushes, PRs require green status checks.
|
|
|
|
|
- All work (human or AI) happens on feature branches → PR to `main`.
|
2026-07-13 08:38:37 +01:00
|
|
|
- Merging to `main` releases **to staging only**. Production moves only on
|
|
|
|
|
the second approval. If staging or the E2E gate fails, fix forward —
|
|
|
|
|
production is untouched either way.
|
|
|
|
|
|
|
|
|
|
## Promotion granularity
|
|
|
|
|
|
|
|
|
|
Staging always runs the latest `main`. Promoting approves a *state of main*,
|
|
|
|
|
not a single PR — if two PRs merged since the last promotion, they ship
|
|
|
|
|
together. Test staging accordingly. To promote an older state, pass its
|
|
|
|
|
commit SHA to the promote workflow (its images must still exist in the
|
|
|
|
|
registry).
|
|
|
|
|
|
|
|
|
|
Staging quirk for manual testing: external `/api` is broken at the staging
|
|
|
|
|
proxy — exercise API endpoints from the host, not via the public staging URL.
|
2026-07-03 06:50:40 +01:00
|
|
|
|
|
|
|
|
## Environments
|
|
|
|
|
|
|
|
|
|
| | Production | Staging |
|
|
|
|
|
|---|---|---|
|
|
|
|
|
| Portainer stack file | `docker-compose.portainer.yml` | `docker-compose.portainer.staging.yml` |
|
|
|
|
|
| Image tag | `:prod` | `:staging` |
|
|
|
|
|
| Container prefix | `sc_` / `schoolcompare_` | `sc_staging_` |
|
|
|
|
|
| Frontend macvlan IP | 10.0.1.150 | `STAGING_FRONTEND_IP` (default 10.0.1.151) |
|
|
|
|
|
| Postgres macvlan IP | 10.0.1.189 | `STAGING_DB_IP` (default 10.0.1.190) |
|
|
|
|
|
| Airflow UI port | 8080 | 8081 |
|
|
|
|
|
| Volumes | stack-prefixed | stack-prefixed (fully isolated) |
|
|
|
|
|
|
|
|
|
|
Staging gets `:staging` images on every merge to main — even ones that later
|
|
|
|
|
fail the E2E gate. That's the point: staging absorbs the risk.
|
|
|
|
|
|
|
|
|
|
## Gitea repository secrets
|
|
|
|
|
|
|
|
|
|
| Secret | Purpose |
|
|
|
|
|
|---|---|
|
2026-07-03 13:39:59 +01:00
|
|
|
| `REGISTRY_TOKEN` | push images to privaterepo.sitaru.org (already set) |
|
2026-07-03 08:28:20 +01:00
|
|
|
| `CLAUDE_CODE_OAUTH_TOKEN` | Claude Code subscription auth for the PR review — generate with `claude setup-token` on your machine |
|
2026-07-03 06:50:40 +01:00
|
|
|
| `PORTAINER_STAGING_WEBHOOK` | staging stack redeploy webhook URL |
|
|
|
|
|
| `PORTAINER_PROD_WEBHOOK` | production stack redeploy webhook URL |
|
|
|
|
|
| `STAGING_BASE_URL` | e.g. `http://10.0.1.151:3000` — health poll + E2E target |
|
|
|
|
|
| `PROD_BASE_URL` | e.g. `http://10.0.1.150:3000` — post-promotion health poll |
|
|
|
|
|
|
|
|
|
|
## One-time setup checklist
|
|
|
|
|
|
|
|
|
|
1. **Create the staging stack** in Portainer from
|
|
|
|
|
`docker-compose.portainer.staging.yml` (stack name e.g.
|
|
|
|
|
`schoolcompare-staging`). Set the same environment variables as prod plus
|
|
|
|
|
`STAGING_DB_IP` / `STAGING_FRONTEND_IP` if the defaults clash.
|
|
|
|
|
2. **Enable webhooks** on both stacks (Portainer → Stack → Webhook) and store
|
|
|
|
|
the URLs as `PORTAINER_STAGING_WEBHOOK` / `PORTAINER_PROD_WEBHOOK`. Remove
|
|
|
|
|
the old hardcoded webhook usage (now gone from the workflows).
|
|
|
|
|
3. **Add the remaining secrets** listed above in Gitea → repo → Settings →
|
|
|
|
|
Actions → Secrets.
|
|
|
|
|
4. **Protect `main`** in Gitea → Settings → Branches: require PRs, require the
|
|
|
|
|
pr-checks status checks (frontend, backend, builds, ai-review) to pass.
|
|
|
|
|
5. **Bootstrap staging data via Airflow** (no prod dump — staging populates
|
|
|
|
|
itself from source, exercising the pipeline image end-to-end):
|
|
|
|
|
- Open the staging Airflow UI (`http://<host>:8081`) and trigger, in order:
|
2026-07-06 09:01:26 +01:00
|
|
|
`school_data_daily`, `school_data_monthly_ofsted`, then the manual-schedule
|
2026-07-03 06:50:40 +01:00
|
|
|
`school_data_annual_ees` and `school_data_annual_idaci`.
|
|
|
|
|
- First runs download from government sources (GIAS, Ofsted, EES, IDACI),
|
|
|
|
|
run dbt, and sync Typesense — expect the initial backfill to take a while.
|
|
|
|
|
- The scheduled DAGs then keep staging fresh exactly like prod.
|
|
|
|
|
6. **Switch the prod stack to `:prod` tags** — the repo's
|
|
|
|
|
`docker-compose.portainer.yml` is already updated; redeploy the prod stack
|
|
|
|
|
from it. Until the first pipeline run promotes an image, tag the current
|
|
|
|
|
images manually: `docker buildx imagetools create -t <image>:prod <image>:latest`
|
|
|
|
|
for each of the three images.
|
|
|
|
|
|
|
|
|
|
## Rollback
|
|
|
|
|
|
2026-07-13 08:38:37 +01:00
|
|
|
Re-run "Promote to Production (manual)" with the SHA of the last good commit
|
|
|
|
|
(fastest, fully gated), or manually re-point the tags — every promotion first
|
|
|
|
|
saves the outgoing `:prod` as `:prod-previous`:
|
2026-07-03 06:50:40 +01:00
|
|
|
|
|
|
|
|
```bash
|
|
|
|
|
for img in backend frontend pipeline; do
|
|
|
|
|
docker buildx imagetools create \
|
|
|
|
|
-t privaterepo.sitaru.org/tudor/school_compare-$img:prod \
|
|
|
|
|
privaterepo.sitaru.org/tudor/school_compare-$img:prod-previous
|
|
|
|
|
done
|
|
|
|
|
curl -fsSk -X POST "$PORTAINER_PROD_WEBHOOK"
|
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
Or promote any older build directly: `imagetools create -t <image>:prod <image>:sha-<shortsha>`.
|
|
|
|
|
|
|
|
|
|
## E2E suite
|
|
|
|
|
|
|
|
|
|
Lives in `e2e/` (own package — CI installs it without the app's node_modules).
|
|
|
|
|
Journeys: home + name search, postcode search, school detail, two-school
|
|
|
|
|
comparison, rankings table. Run locally against any environment:
|
|
|
|
|
|
|
|
|
|
```bash
|
|
|
|
|
cd e2e && npm ci
|
|
|
|
|
BASE_URL=http://10.0.1.151:3000 npx playwright test
|
|
|
|
|
```
|
|
|
|
|
|
|
|
|
|
Tests assert data invariants (results exist, charts render), not exact
|
|
|
|
|
numbers, so scheduled data refreshes don't break the gate.
|
|
|
|
|
|
|
|
|
|
## AI code review
|
|
|
|
|
|
2026-07-03 08:28:20 +01:00
|
|
|
`scripts/ci/ai_review.py` pipes the PR diff through headless Claude Code
|
|
|
|
|
(`claude -p`, authenticated with the subscription OAuth token — no API
|
2026-07-03 13:39:59 +01:00
|
|
|
billing), posts the structured findings as a PR comment using the per-run
|
|
|
|
|
token Gitea Actions provides automatically (`secrets.GITEA_TOKEN` — no setup
|
|
|
|
|
needed), and fails the check only when a finding is rated
|
2026-07-03 08:28:20 +01:00
|
|
|
**severe** (would break prod, leak data, or corrupt data). Minor findings are
|
|
|
|
|
informational and never block a merge.
|
2026-08-23 10:51:35 +01:00
|
|
|
|
2026-08-26 20:56:32 +01:00
|
|
|
## Rate limiting, and the Cloudflare gap
|
|
|
|
|
|
|
|
|
|
Two independent limits protect the API:
|
|
|
|
|
|
|
|
|
|
- **Per client**, via slowapi, keyed on `CF-Connecting-IP` (falling back to
|
|
|
|
|
`X-Forwarded-For`, then the peer address). 60/minute by default;
|
|
|
|
|
`/api/suggest` gets 120/minute because typing is bursty.
|
|
|
|
|
- **Globally**, via `GlobalRateLimitMiddleware`: a fixed 60-second window over
|
|
|
|
|
all `/api/` traffic, `GLOBAL_RATE_LIMIT_PER_MINUTE` (default 3000),
|
|
|
|
|
independent of any client identity. Requests from `127.0.0.1` are exempt so
|
|
|
|
|
the container healthcheck cannot be starved into a restart loop.
|
|
|
|
|
|
|
|
|
|
### Open: the origin must only accept Cloudflare
|
|
|
|
|
|
|
|
|
|
`CF-Connecting-IP` is only meaningful for requests that actually reached the
|
|
|
|
|
origin through Cloudflare, and **the application cannot verify that they did**.
|
|
|
|
|
Anything able to reach the origin directly can set that header freely and, by
|
|
|
|
|
rotating it, mint a fresh rate-limit bucket per request — defeating per-client
|
|
|
|
|
limits on every endpoint.
|
|
|
|
|
|
|
|
|
|
The global ceiling bounds the damage to total origin capacity. It does not fix
|
|
|
|
|
the underlying gap, and nothing in the code can. Closing it needs one of:
|
|
|
|
|
|
|
|
|
|
- **Authenticated Origin Pulls** — Cloudflare presents a client certificate the
|
|
|
|
|
origin requires, so non-Cloudflare traffic is refused at TLS.
|
|
|
|
|
- **An origin firewall** restricted to Cloudflare's published IP ranges.
|
|
|
|
|
|
|
|
|
|
Until one is in place, treat per-client limits as protection against accidents
|
|
|
|
|
and ordinary load, not against a determined caller.
|
|
|
|
|
|
2026-08-23 10:51:35 +01:00
|
|
|
## Feature flags (Unleash)
|
|
|
|
|
|
|
|
|
|
Flag state lives in a self-hosted Unleash instance, deployed as its own
|
|
|
|
|
Portainer stack from `docker-compose.portainer.unleash.yml`. It is separate
|
|
|
|
|
from the application stacks on purpose — redeploying staging must not be able
|
|
|
|
|
to disturb production's flags.
|
|
|
|
|
|
|
|
|
|
The flags themselves are declared in `backend/flags.py`. Unleash holds the
|
|
|
|
|
state; the registry holds the list. A flag in the UI that is not in the
|
|
|
|
|
registry is orphaned and nothing reads it.
|
|
|
|
|
|
|
|
|
|
### First-time setup
|
|
|
|
|
|
|
|
|
|
1. Deploy the stack in Portainer. Set `UNLEASH_DB_PASSWORD`,
|
|
|
|
|
`UNLEASH_ADMIN_PASSWORD` and (optionally) `UNLEASH_IP`.
|
|
|
|
|
2. Log in to the UI at `http://<UNLEASH_IP>:4242` as `admin`.
|
|
|
|
|
3. Create one **client** API token per environment:
|
|
|
|
|
- `schoolcompare-staging`, environment **development**
|
|
|
|
|
- `schoolcompare-prod`, environment **production**
|
|
|
|
|
|
|
|
|
|
Client tokens, not admin tokens — the backend only reads.
|
|
|
|
|
4. Put each token in the matching Portainer stack's `UNLEASH_API_TOKEN`
|
|
|
|
|
variable, and set `UNLEASH_URL` to `http://<UNLEASH_IP>:4242/api`.
|
|
|
|
|
5. Redeploy the application stacks.
|
|
|
|
|
|
|
|
|
|
### Turning a feature on
|
|
|
|
|
|
|
|
|
|
Toggle the flag in the environment you want. Flags appear in the Unleash UI
|
|
|
|
|
after the backend has evaluated them once, so a newly declared flag shows up
|
|
|
|
|
shortly after the deploy that introduced it.
|
|
|
|
|
|
|
|
|
|
A flip reaches school pages within about five minutes and place pages within
|
|
|
|
|
the hour. Next's ISR does the propagating — it revalidates a route at the
|
|
|
|
|
*lowest* `revalidate` among that route's fetches, which is 300s for
|
|
|
|
|
`/school/[slug]` and 3600s for the place pages. There is no webhook, and
|
|
|
|
|
adding one would only be worth it if flips ever needed to be instant.
|
|
|
|
|
|
|
|
|
|
### When Unleash is unreachable
|
|
|
|
|
|
|
|
|
|
Every flag evaluates to `False` and the site serves as though nothing were
|
|
|
|
|
switched on. That is deliberate — an unfinished feature staying hidden is the
|
|
|
|
|
safe direction — but it means a *released* feature disappears if a backend
|
|
|
|
|
container cold-starts with an empty cache while Unleash is down. The SDK's
|
|
|
|
|
disk cache is on a named volume so restarts keep last-known state, and flags
|
|
|
|
|
are removed from the code within 90 days (enforced by a test), which bounds
|
|
|
|
|
how long any feature is exposed to this.
|
|
|
|
|
|
|
|
|
|
If `UNLEASH_URL` is unset, every flag is `False` and no connection is
|
|
|
|
|
attempted. That is the correct behaviour for local development and CI, and it
|
|
|
|
|
means the test suites need no flag server.
|