The runbook said a flag 'appears in the Unleash UI after the backend has evaluated it once'. That is wrong. SDKs read definitions from the server and never register anything, and metrics for an unknown flag are discarded — so a declared flag is evaluated on every request, stays False forever, and never shows up until someone creates it by hand. Found the way these things usually are: staging had been running the flag code for a while and the UI was still empty. Also names the environment trap while here — each stack's token is scoped to one environment, so toggling the other does nothing visible and looks like the flag is broken. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
10 KiB
SDLC & Deployment Pipeline
SchoolCompare uses a two-stage deploy model on Gitea Actions with two human approvals. AI writes the code on feature branches; the first approval merges the PR, which deploys to staging and runs the E2E gate; the second approval — after manual testing on staging — promotes the exact same images to production via a manual workflow.
The flow
feature branch (AI-authored)
│ PR to main ← approval #1
▼
PR checks (.gitea/workflows/pr-checks.yml)
typecheck + unit tests + backend smoke + image builds (no push)
+ Claude code review posted as a PR comment (severe findings fail the check)
│ merge (branch protection requires green checks)
▼
Stage pipeline (.gitea/workflows/deploy.yml) — automatic
1. build & push images → tags sha-<sha>, staging
2. staging Portainer webhook → wait for staging health
3. Playwright E2E journeys against staging ← gate before human testing
▼
Manual testing on staging (stx.schoolcompare.co.uk)
│ Actions → "Promote to Production (manual)" ← approval #2
▼
Promote pipeline (.gitea/workflows/promote.yml) — manual dispatch
1. resolve target sha (input, or latest main if empty)
2. REFUSE unless that commit's "E2E Journeys against Staging" status is green
3. retag sha-<sha> → :prod (same bytes — build once, promote the image)
previous :prod saved as :prod-previous
4. prod Portainer webhook → wait for prod health
Key principle: build once, promote the exact image. Production pins :prod,
which only moves when a human runs the promote workflow — and the workflow
only accepts commits that passed the staging E2E gate. Nothing tags :latest
anymore.
Branch & PR workflow
mainis protected: no direct pushes, PRs require green status checks.- All work (human or AI) happens on feature branches → PR to
main. - Merging to
mainreleases to staging only. Production moves only on the second approval. If staging or the E2E gate fails, fix forward — production is untouched either way.
Promotion granularity
Staging always runs the latest main. Promoting approves a state of main,
not a single PR — if two PRs merged since the last promotion, they ship
together. Test staging accordingly. To promote an older state, pass its
commit SHA to the promote workflow (its images must still exist in the
registry).
Staging quirk for manual testing: external /api is broken at the staging
proxy — exercise API endpoints from the host, not via the public staging URL.
Environments
| Production | Staging | |
|---|---|---|
| Portainer stack file | docker-compose.portainer.yml |
docker-compose.portainer.staging.yml |
| Image tag | :prod |
:staging |
| Container prefix | sc_ / schoolcompare_ |
sc_staging_ |
| Frontend macvlan IP | 10.0.1.150 | STAGING_FRONTEND_IP (default 10.0.1.151) |
| Postgres macvlan IP | 10.0.1.189 | STAGING_DB_IP (default 10.0.1.190) |
| Airflow UI port | 8080 | 8081 |
| Volumes | stack-prefixed | stack-prefixed (fully isolated) |
Staging gets :staging images on every merge to main — even ones that later
fail the E2E gate. That's the point: staging absorbs the risk.
Gitea repository secrets
| Secret | Purpose |
|---|---|
REGISTRY_TOKEN |
push images to privaterepo.sitaru.org (already set) |
CLAUDE_CODE_OAUTH_TOKEN |
Claude Code subscription auth for the PR review — generate with claude setup-token on your machine |
PORTAINER_STAGING_WEBHOOK |
staging stack redeploy webhook URL |
PORTAINER_PROD_WEBHOOK |
production stack redeploy webhook URL |
STAGING_BASE_URL |
e.g. http://10.0.1.151:3000 — health poll + E2E target |
PROD_BASE_URL |
e.g. http://10.0.1.150:3000 — post-promotion health poll |
One-time setup checklist
- Create the staging stack in Portainer from
docker-compose.portainer.staging.yml(stack name e.g.schoolcompare-staging). Set the same environment variables as prod plusSTAGING_DB_IP/STAGING_FRONTEND_IPif the defaults clash. - Enable webhooks on both stacks (Portainer → Stack → Webhook) and store
the URLs as
PORTAINER_STAGING_WEBHOOK/PORTAINER_PROD_WEBHOOK. Remove the old hardcoded webhook usage (now gone from the workflows). - Add the remaining secrets listed above in Gitea → repo → Settings → Actions → Secrets.
- Protect
mainin Gitea → Settings → Branches: require PRs, require the pr-checks status checks (frontend, backend, builds, ai-review) to pass. - Bootstrap staging data via Airflow (no prod dump — staging populates
itself from source, exercising the pipeline image end-to-end):
- Open the staging Airflow UI (
http://<host>:8081) and trigger, in order:school_data_daily,school_data_monthly_ofsted, then the manual-scheduleschool_data_annual_eesandschool_data_annual_idaci. - First runs download from government sources (GIAS, Ofsted, EES, IDACI), run dbt, and sync Typesense — expect the initial backfill to take a while.
- The scheduled DAGs then keep staging fresh exactly like prod.
- Open the staging Airflow UI (
- Switch the prod stack to
:prodtags — the repo'sdocker-compose.portainer.ymlis already updated; redeploy the prod stack from it. Until the first pipeline run promotes an image, tag the current images manually:docker buildx imagetools create -t <image>:prod <image>:latestfor each of the three images.
Rollback
Re-run "Promote to Production (manual)" with the SHA of the last good commit
(fastest, fully gated), or manually re-point the tags — every promotion first
saves the outgoing :prod as :prod-previous:
for img in backend frontend pipeline; do
docker buildx imagetools create \
-t privaterepo.sitaru.org/tudor/school_compare-$img:prod \
privaterepo.sitaru.org/tudor/school_compare-$img:prod-previous
done
curl -fsSk -X POST "$PORTAINER_PROD_WEBHOOK"
Or promote any older build directly: imagetools create -t <image>:prod <image>:sha-<shortsha>.
E2E suite
Lives in e2e/ (own package — CI installs it without the app's node_modules).
Journeys: home + name search, postcode search, school detail, two-school
comparison, rankings table. Run locally against any environment:
cd e2e && npm ci
BASE_URL=http://10.0.1.151:3000 npx playwright test
Tests assert data invariants (results exist, charts render), not exact numbers, so scheduled data refreshes don't break the gate.
AI code review
scripts/ci/ai_review.py pipes the PR diff through headless Claude Code
(claude -p, authenticated with the subscription OAuth token — no API
billing), posts the structured findings as a PR comment using the per-run
token Gitea Actions provides automatically (secrets.GITEA_TOKEN — no setup
needed), and fails the check only when a finding is rated
severe (would break prod, leak data, or corrupt data). Minor findings are
informational and never block a merge.
Feature flags (Unleash)
Flag state lives in a self-hosted Unleash instance, deployed as its own
Portainer stack from docker-compose.portainer.unleash.yml. It is separate
from the application stacks on purpose — redeploying staging must not be able
to disturb production's flags.
The flags themselves are declared in backend/flags.py. Unleash holds the
state; the registry holds the list. A flag in the UI that is not in the
registry is orphaned and nothing reads it.
First-time setup
-
Deploy the stack in Portainer. Set
UNLEASH_DB_PASSWORD,UNLEASH_ADMIN_PASSWORDand (optionally)UNLEASH_IP. -
Log in to the UI at
http://<UNLEASH_IP>:4242asadmin. -
Create one client API token per environment:
schoolcompare-staging, environment developmentschoolcompare-prod, environment production
Client tokens, not admin tokens — the backend only reads.
-
Put each token in the matching Portainer stack's
UNLEASH_API_TOKENvariable, and setUNLEASH_URLtohttp://<UNLEASH_IP>:4242/api. -
Redeploy the application stacks.
Adding a flag to Unleash
Unleash does not create flags by itself. The SDK reads definitions from the
server and never registers anything, and metrics for a flag the server has
never heard of are discarded. So a flag declared in backend/flags.py will be
evaluated on every request, stay False forever, and never appear in the UI
until someone creates it there by hand.
For each flag in the registry, create one in Unleash with:
- Name — character for character what
backend/flags.pydeclares. snake_case, no hyphens or spaces. A typo produces a flag that looks correct in the UI and is read by nothing. - Type — Release. No strategies, constraints or variants: these are plain on/off switches, by design.
Turning a feature on
Toggle the flag in the environment matching the stack you mean: development for staging, production for prod. The token in each stack is scoped to one environment, so toggling the other one has no visible effect.
The SDK refreshes every 15 seconds, so the API reflects the change almost at once; the pages follow on their own schedule, below.
A flip reaches school pages within about five minutes and place pages within
the hour. Next's ISR does the propagating — it revalidates a route at the
lowest revalidate among that route's fetches, which is 300s for
/school/[slug] and 3600s for the place pages. There is no webhook, and
adding one would only be worth it if flips ever needed to be instant.
When Unleash is unreachable
Every flag evaluates to False and the site serves as though nothing were
switched on. That is deliberate — an unfinished feature staying hidden is the
safe direction — but it means a released feature disappears if a backend
container cold-starts with an empty cache while Unleash is down. The SDK's
disk cache is on a named volume so restarts keep last-known state, and flags
are removed from the code within 90 days (enforced by a test), which bounds
how long any feature is exposed to this.
If UNLEASH_URL is unset, every flag is False and no connection is
attempted. That is the correct behaviour for local development and CI, and it
means the test suites need no flag server.