PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 18s
PR Checks / Build Frontend (no push) (pull_request) Successful in 45s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 10s
Code review, both findings valid. The design doc claimed Cloudflare "replaces the header, so a browser cannot forge it", and that only the X-Forwarded-For fallback was forgeable. That is true only for traffic that actually passed through Cloudflare, and nothing in this process can verify that it did. Reaching the origin directly, both headers are equally attacker-controlled — and rotating CF-Connecting-IP mints a fresh rate-limit bucket per request, defeating per-client limits on every endpoint including the DataFrame-heavy /api/schools. Against abuse that is worse than the shared bucket it replaced, which at least capped everyone together. So the ceiling comes back. I dropped it earlier arguing it belonged at Cloudflare; that argument assumed the keying was sound, and it is not. GlobalRateLimitMiddleware counts all /api/ traffic in a fixed window against a total, independent of client identity, outermost so it refuses before any work happens. Written by hand because slowapi cannot express a global cap: default_limits and application_limits are both keyed by key_func, and the latter needs middleware this app does not install. It does not make the header trustworthy — it makes trusting it survivable. The real fix is Authenticated Origin Pulls or an origin firewall, now documented in DEPLOY.md as the open gap it is. 127.0.0.1 is exempt: the healthcheck curls localhost from inside the container, and starving it would restart the container and turn a load spike into an outage loop. Keyed on the peer address, never the Host header, which the caller sets. Second finding: suggest_schools_typesense promised "never raises" while the parsing loop sat outside the try, so int(None) on a malformed document would have made a keystroke a 500. The loop now skips bad rows rather than dropping the whole list — and a hit with no document no longer becomes a suggestion pointing at /school/0. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
234 lines
11 KiB
Markdown
234 lines
11 KiB
Markdown
# SDLC & Deployment Pipeline
|
|
|
|
SchoolCompare uses a two-stage deploy model on Gitea Actions with two human
|
|
approvals. AI writes the code on feature branches; the first approval merges
|
|
the PR, which deploys to staging and runs the E2E gate; the second approval —
|
|
after manual testing on staging — promotes the exact same images to
|
|
production via a manual workflow.
|
|
|
|
## The flow
|
|
|
|
```
|
|
feature branch (AI-authored)
|
|
│ PR to main ← approval #1
|
|
▼
|
|
PR checks (.gitea/workflows/pr-checks.yml)
|
|
typecheck + unit tests + backend smoke + image builds (no push)
|
|
+ Claude code review posted as a PR comment (severe findings fail the check)
|
|
│ merge (branch protection requires green checks)
|
|
▼
|
|
Stage pipeline (.gitea/workflows/deploy.yml) — automatic
|
|
1. build & push images → tags sha-<sha>, staging
|
|
2. staging Portainer webhook → wait for staging health
|
|
3. Playwright E2E journeys against staging ← gate before human testing
|
|
▼
|
|
Manual testing on staging (stx.schoolcompare.co.uk)
|
|
│ Actions → "Promote to Production (manual)" ← approval #2
|
|
▼
|
|
Promote pipeline (.gitea/workflows/promote.yml) — manual dispatch
|
|
1. resolve target sha (input, or latest main if empty)
|
|
2. REFUSE unless that commit's "E2E Journeys against Staging" status is green
|
|
3. retag sha-<sha> → :prod (same bytes — build once, promote the image)
|
|
previous :prod saved as :prod-previous
|
|
4. prod Portainer webhook → wait for prod health
|
|
```
|
|
|
|
Key principle: **build once, promote the exact image**. Production pins `:prod`,
|
|
which only moves when a human runs the promote workflow — and the workflow
|
|
only accepts commits that passed the staging E2E gate. Nothing tags `:latest`
|
|
anymore.
|
|
|
|
## Branch & PR workflow
|
|
|
|
- `main` is protected: no direct pushes, PRs require green status checks.
|
|
- All work (human or AI) happens on feature branches → PR to `main`.
|
|
- Merging to `main` releases **to staging only**. Production moves only on
|
|
the second approval. If staging or the E2E gate fails, fix forward —
|
|
production is untouched either way.
|
|
|
|
## Promotion granularity
|
|
|
|
Staging always runs the latest `main`. Promoting approves a *state of main*,
|
|
not a single PR — if two PRs merged since the last promotion, they ship
|
|
together. Test staging accordingly. To promote an older state, pass its
|
|
commit SHA to the promote workflow (its images must still exist in the
|
|
registry).
|
|
|
|
Staging quirk for manual testing: external `/api` is broken at the staging
|
|
proxy — exercise API endpoints from the host, not via the public staging URL.
|
|
|
|
## Environments
|
|
|
|
| | Production | Staging |
|
|
|---|---|---|
|
|
| Portainer stack file | `docker-compose.portainer.yml` | `docker-compose.portainer.staging.yml` |
|
|
| Image tag | `:prod` | `:staging` |
|
|
| Container prefix | `sc_` / `schoolcompare_` | `sc_staging_` |
|
|
| Frontend macvlan IP | 10.0.1.150 | `STAGING_FRONTEND_IP` (default 10.0.1.151) |
|
|
| Postgres macvlan IP | 10.0.1.189 | `STAGING_DB_IP` (default 10.0.1.190) |
|
|
| Airflow UI port | 8080 | 8081 |
|
|
| Volumes | stack-prefixed | stack-prefixed (fully isolated) |
|
|
|
|
Staging gets `:staging` images on every merge to main — even ones that later
|
|
fail the E2E gate. That's the point: staging absorbs the risk.
|
|
|
|
## Gitea repository secrets
|
|
|
|
| Secret | Purpose |
|
|
|---|---|
|
|
| `REGISTRY_TOKEN` | push images to privaterepo.sitaru.org (already set) |
|
|
| `CLAUDE_CODE_OAUTH_TOKEN` | Claude Code subscription auth for the PR review — generate with `claude setup-token` on your machine |
|
|
| `PORTAINER_STAGING_WEBHOOK` | staging stack redeploy webhook URL |
|
|
| `PORTAINER_PROD_WEBHOOK` | production stack redeploy webhook URL |
|
|
| `STAGING_BASE_URL` | e.g. `http://10.0.1.151:3000` — health poll + E2E target |
|
|
| `PROD_BASE_URL` | e.g. `http://10.0.1.150:3000` — post-promotion health poll |
|
|
|
|
## One-time setup checklist
|
|
|
|
1. **Create the staging stack** in Portainer from
|
|
`docker-compose.portainer.staging.yml` (stack name e.g.
|
|
`schoolcompare-staging`). Set the same environment variables as prod plus
|
|
`STAGING_DB_IP` / `STAGING_FRONTEND_IP` if the defaults clash.
|
|
2. **Enable webhooks** on both stacks (Portainer → Stack → Webhook) and store
|
|
the URLs as `PORTAINER_STAGING_WEBHOOK` / `PORTAINER_PROD_WEBHOOK`. Remove
|
|
the old hardcoded webhook usage (now gone from the workflows).
|
|
3. **Add the remaining secrets** listed above in Gitea → repo → Settings →
|
|
Actions → Secrets.
|
|
4. **Protect `main`** in Gitea → Settings → Branches: require PRs, require the
|
|
pr-checks status checks (frontend, backend, builds, ai-review) to pass.
|
|
5. **Bootstrap staging data via Airflow** (no prod dump — staging populates
|
|
itself from source, exercising the pipeline image end-to-end):
|
|
- Open the staging Airflow UI (`http://<host>:8081`) and trigger, in order:
|
|
`school_data_daily`, `school_data_monthly_ofsted`, then the manual-schedule
|
|
`school_data_annual_ees` and `school_data_annual_idaci`.
|
|
- First runs download from government sources (GIAS, Ofsted, EES, IDACI),
|
|
run dbt, and sync Typesense — expect the initial backfill to take a while.
|
|
- The scheduled DAGs then keep staging fresh exactly like prod.
|
|
6. **Switch the prod stack to `:prod` tags** — the repo's
|
|
`docker-compose.portainer.yml` is already updated; redeploy the prod stack
|
|
from it. Until the first pipeline run promotes an image, tag the current
|
|
images manually: `docker buildx imagetools create -t <image>:prod <image>:latest`
|
|
for each of the three images.
|
|
|
|
## Rollback
|
|
|
|
Re-run "Promote to Production (manual)" with the SHA of the last good commit
|
|
(fastest, fully gated), or manually re-point the tags — every promotion first
|
|
saves the outgoing `:prod` as `:prod-previous`:
|
|
|
|
```bash
|
|
for img in backend frontend pipeline; do
|
|
docker buildx imagetools create \
|
|
-t privaterepo.sitaru.org/tudor/school_compare-$img:prod \
|
|
privaterepo.sitaru.org/tudor/school_compare-$img:prod-previous
|
|
done
|
|
curl -fsSk -X POST "$PORTAINER_PROD_WEBHOOK"
|
|
```
|
|
|
|
Or promote any older build directly: `imagetools create -t <image>:prod <image>:sha-<shortsha>`.
|
|
|
|
## E2E suite
|
|
|
|
Lives in `e2e/` (own package — CI installs it without the app's node_modules).
|
|
Journeys: home + name search, postcode search, school detail, two-school
|
|
comparison, rankings table. Run locally against any environment:
|
|
|
|
```bash
|
|
cd e2e && npm ci
|
|
BASE_URL=http://10.0.1.151:3000 npx playwright test
|
|
```
|
|
|
|
Tests assert data invariants (results exist, charts render), not exact
|
|
numbers, so scheduled data refreshes don't break the gate.
|
|
|
|
## AI code review
|
|
|
|
`scripts/ci/ai_review.py` pipes the PR diff through headless Claude Code
|
|
(`claude -p`, authenticated with the subscription OAuth token — no API
|
|
billing), posts the structured findings as a PR comment using the per-run
|
|
token Gitea Actions provides automatically (`secrets.GITEA_TOKEN` — no setup
|
|
needed), and fails the check only when a finding is rated
|
|
**severe** (would break prod, leak data, or corrupt data). Minor findings are
|
|
informational and never block a merge.
|
|
|
|
## Rate limiting, and the Cloudflare gap
|
|
|
|
Two independent limits protect the API:
|
|
|
|
- **Per client**, via slowapi, keyed on `CF-Connecting-IP` (falling back to
|
|
`X-Forwarded-For`, then the peer address). 60/minute by default;
|
|
`/api/suggest` gets 120/minute because typing is bursty.
|
|
- **Globally**, via `GlobalRateLimitMiddleware`: a fixed 60-second window over
|
|
all `/api/` traffic, `GLOBAL_RATE_LIMIT_PER_MINUTE` (default 3000),
|
|
independent of any client identity. Requests from `127.0.0.1` are exempt so
|
|
the container healthcheck cannot be starved into a restart loop.
|
|
|
|
### Open: the origin must only accept Cloudflare
|
|
|
|
`CF-Connecting-IP` is only meaningful for requests that actually reached the
|
|
origin through Cloudflare, and **the application cannot verify that they did**.
|
|
Anything able to reach the origin directly can set that header freely and, by
|
|
rotating it, mint a fresh rate-limit bucket per request — defeating per-client
|
|
limits on every endpoint.
|
|
|
|
The global ceiling bounds the damage to total origin capacity. It does not fix
|
|
the underlying gap, and nothing in the code can. Closing it needs one of:
|
|
|
|
- **Authenticated Origin Pulls** — Cloudflare presents a client certificate the
|
|
origin requires, so non-Cloudflare traffic is refused at TLS.
|
|
- **An origin firewall** restricted to Cloudflare's published IP ranges.
|
|
|
|
Until one is in place, treat per-client limits as protection against accidents
|
|
and ordinary load, not against a determined caller.
|
|
|
|
## Feature flags (Unleash)
|
|
|
|
Flag state lives in a self-hosted Unleash instance, deployed as its own
|
|
Portainer stack from `docker-compose.portainer.unleash.yml`. It is separate
|
|
from the application stacks on purpose — redeploying staging must not be able
|
|
to disturb production's flags.
|
|
|
|
The flags themselves are declared in `backend/flags.py`. Unleash holds the
|
|
state; the registry holds the list. A flag in the UI that is not in the
|
|
registry is orphaned and nothing reads it.
|
|
|
|
### First-time setup
|
|
|
|
1. Deploy the stack in Portainer. Set `UNLEASH_DB_PASSWORD`,
|
|
`UNLEASH_ADMIN_PASSWORD` and (optionally) `UNLEASH_IP`.
|
|
2. Log in to the UI at `http://<UNLEASH_IP>:4242` as `admin`.
|
|
3. Create one **client** API token per environment:
|
|
- `schoolcompare-staging`, environment **development**
|
|
- `schoolcompare-prod`, environment **production**
|
|
|
|
Client tokens, not admin tokens — the backend only reads.
|
|
4. Put each token in the matching Portainer stack's `UNLEASH_API_TOKEN`
|
|
variable, and set `UNLEASH_URL` to `http://<UNLEASH_IP>:4242/api`.
|
|
5. Redeploy the application stacks.
|
|
|
|
### Turning a feature on
|
|
|
|
Toggle the flag in the environment you want. Flags appear in the Unleash UI
|
|
after the backend has evaluated them once, so a newly declared flag shows up
|
|
shortly after the deploy that introduced it.
|
|
|
|
A flip reaches school pages within about five minutes and place pages within
|
|
the hour. Next's ISR does the propagating — it revalidates a route at the
|
|
*lowest* `revalidate` among that route's fetches, which is 300s for
|
|
`/school/[slug]` and 3600s for the place pages. There is no webhook, and
|
|
adding one would only be worth it if flips ever needed to be instant.
|
|
|
|
### When Unleash is unreachable
|
|
|
|
Every flag evaluates to `False` and the site serves as though nothing were
|
|
switched on. That is deliberate — an unfinished feature staying hidden is the
|
|
safe direction — but it means a *released* feature disappears if a backend
|
|
container cold-starts with an empty cache while Unleash is down. The SDK's
|
|
disk cache is on a named volume so restarts keep last-known state, and flags
|
|
are removed from the code within 90 days (enforced by a test), which bounds
|
|
how long any feature is exposed to this.
|
|
|
|
If `UNLEASH_URL` is unset, every flag is `False` and no connection is
|
|
attempted. That is the correct behaviour for local development and CI, and it
|
|
means the test suites need no flag server.
|