Commit Graph
11 Commits
Author SHA1 Message Date
TudorandClaude Opus 5 0c901cd0d1 feat(ci): gate promotion on the image set that actually passed E2E
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m11s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 18s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m19s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 36s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 6m3s
Staging health polling asked only whether something answered HTTP 200 at
the base URL. It could not tell the new deployment from the old one, so
journeys could pass against the previous release, and concurrent merges
could move the staging tags underneath a run in flight.

Each staging run now mints a build ID and stamps all three images with
the commit and that ID, as labels and — for frontend and backend — as a
build-time JSON file that environment overrides cannot rewrite.
/release.json reports both identities uncached, and scripts/ci/release.py
polls for the expected pair before and after the journeys. Only then are
the captured build digests tagged verified-<sha>.

Promotion resolves those verified tags to immutable digests, revalidates
their labels, and refuses a mixed or incomplete set before any :prod tag
moves. The whole staging workflow shares one concurrency group with
cancellation disabled, so releases serialise.

The scripts are stdlib-only and unit-tested against mocked registry and
HTTP behaviour; PR checks now run the pipeline and CI suites too. The
runbook records what this cannot prove locally, and that the first
rollout needs a commit built by this workflow.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 10:17:50 +01:00
TudorandClaude Opus 5 264edd2e3a fix(airflow): a login that survives a container restart
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m7s
PR Checks / Backend Smoke (pull_request) Successful in 10s
PR Checks / Build Backend (no push) (pull_request) Successful in 12s
PR Checks / Build Frontend (no push) (pull_request) Successful in 50s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 51s
PR Checks / AI Code Review (Claude) (pull_request) Failing after 2m58s
The simple auth manager generates a random password on first start and
writes it to a file, so every restart of the api-server invalidated the
last one and the password had to be dug out of the container logs again.

The stack now writes that file itself from AIRFLOW_ADMIN_PASSWORD before
exec'ing the api-server. Airflow generates nothing when the file already
exists, so the login is whatever the stack environment says it is.

Written with python rather than echo, so json.dumps escapes a password
containing quotes, backslashes or non-ASCII correctly — verified against
`p@ss "wo\rd' £5`, which round-trips intact.

An unset AIRFLOW_ADMIN_PASSWORD raises KeyError and the container exits.
Falling back to a generated password would silently undo the point of the
change, and a compose-level `:?` gives the same refusal a readable reason.
This does mean the variable MUST be set in Portainer before the next
deploy of either stack.

Not affected by the two Docker gotchas in the upstream docs: this image
has no USER directive so it runs as root, and the file is rewritten from
the environment on every start rather than persisted on a volume.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvdDKvFFSZuMVDH5fEyTob
2026-08-31 17:38:42 +01:00
tudor d8ccb5b733 Merge pull request 'feat(suggest): school autosuggest, and the rate-limit fix it needed first' (#127) from feat/school-autosuggest into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 20s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 51s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 5m49s
Reviewed-on: #127
2026-08-26 19:58:57 +00:00
TudorandClaude Opus 5 0fa1a292c7 fix(api): bound what a forged CF-Connecting-IP can buy
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 18s
PR Checks / Build Frontend (no push) (pull_request) Successful in 45s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 10s
Code review, both findings valid.

The design doc claimed Cloudflare "replaces the header, so a browser
cannot forge it", and that only the X-Forwarded-For fallback was
forgeable. That is true only for traffic that actually passed through
Cloudflare, and nothing in this process can verify that it did. Reaching
the origin directly, both headers are equally attacker-controlled — and
rotating CF-Connecting-IP mints a fresh rate-limit bucket per request,
defeating per-client limits on every endpoint including the
DataFrame-heavy /api/schools. Against abuse that is worse than the
shared bucket it replaced, which at least capped everyone together.

So the ceiling comes back. I dropped it earlier arguing it belonged at
Cloudflare; that argument assumed the keying was sound, and it is not.
GlobalRateLimitMiddleware counts all /api/ traffic in a fixed window
against a total, independent of client identity, outermost so it refuses
before any work happens. Written by hand because slowapi cannot express
a global cap: default_limits and application_limits are both keyed by
key_func, and the latter needs middleware this app does not install.

It does not make the header trustworthy — it makes trusting it
survivable. The real fix is Authenticated Origin Pulls or an origin
firewall, now documented in DEPLOY.md as the open gap it is.

127.0.0.1 is exempt: the healthcheck curls localhost from inside the
container, and starving it would restart the container and turn a load
spike into an outage loop. Keyed on the peer address, never the Host
header, which the caller sets.

Second finding: suggest_schools_typesense promised "never raises" while
the parsing loop sat outside the try, so int(None) on a malformed
document would have made a keystroke a 500. The loop now skips bad rows
rather than dropping the whole list — and a hit with no document no
longer becomes a suggestion pointing at /school/0.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-26 20:56:32 +01:00
TudorandClaude Opus 5 59265f78b6 docs(flags): Unleash does not create flags by itself
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m6s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 35s
PR Checks / Build Frontend (no push) (pull_request) Successful in 46s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 1m15s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 35s
The runbook said a flag 'appears in the Unleash UI after the backend has
evaluated it once'. That is wrong. SDKs read definitions from the server
and never register anything, and metrics for an unknown flag are
discarded — so a declared flag is evaluated on every request, stays
False forever, and never shows up until someone creates it by hand.

Found the way these things usually are: staging had been running the
flag code for a while and the UI was still empty.

Also names the environment trap while here — each stack's token is
scoped to one environment, so toggling the other does nothing visible
and looks like the flag is broken.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-26 20:11:45 +01:00
TudorandClaude Opus 5 01ccbb8e82 feat(flags): add the Unleash stack and its runbook
Its own Portainer stack, belonging to neither application stack: a
staging redeploy must not be able to disturb production's flag state.

One instance serves both. OSS Unleash ships development and production
environments with environment-scoped client tokens, so the same flag
holds independent state in each — which is what lets a feature be on in
staging, where the E2E journeys exercise it, while production stays dark.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-23 10:51:35 +01:00
TudorandClaude Fable 5 2b563cc0bf docs: two-stage deploy model (staging auto, production manual)
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 9m36s
PR Checks / Backend Smoke (pull_request) Successful in 7s
PR Checks / Build Backend (no push) (pull_request) Successful in 10s
PR Checks / Build Frontend (no push) (pull_request) Successful in 43s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Failing after 2m10s
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0146VHeLAWjDVE2B5uU67jCB
2026-07-13 08:38:37 +01:00
TudorandClaude Fable 5 95081d38bd chore: remove the Ofsted Parent View feature end to end
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 9m39s
PR Checks / Backend Smoke (pull_request) Successful in 5s
PR Checks / Build Backend (no push) (pull_request) Successful in 17s
PR Checks / Build Frontend (no push) (pull_request) Successful in 41s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 45s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 1m35s
Removes the 'What Parents Say' section and all supporting elements:

Frontend:
- Drop the OfstedParentView type, the parent_view field, the survey
  section and the 'X% would recommend' callouts in the primary and
  secondary detail views, the Parents nav item, and the parent-view CSS.

Backend:
- Remove the FactParentView model, its loading in data_loader, and
  parent_view from the school-details API response.
- Bump SCHEMA_VERSION to 6 and add an idempotent drop step
  (DROP TABLE IF EXISTS marts.fact_parent_view) to the CLI migration;
  add scripts/sql/drop_fact_parent_view.sql to apply directly to the
  dbt-owned marts DBs on staging and prod.

Pipeline:
- Delete the stg_parent_view + fact_parent_view dbt models and their
  source/schema entries, the tap-uk-parent-view Meltano extractor, and
  the monthly Parent View DAG; drop it from the Dockerfile and the
  staging bootstrap docs.

The rest of dbt (which builds every mart the app reads) is untouched.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-06 09:01:26 +01:00
TudorandClaude Fable 5 c62ba0ca25 fix(ci): post AI review comments with the run-scoped Gitea token
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 9m37s
PR Checks / Backend Smoke (pull_request) Successful in 5s
PR Checks / Build Backend (no push) (pull_request) Successful in 10s
PR Checks / Build Frontend (no push) (pull_request) Successful in 45s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Failing after 1m56s
REGISTRY_TOKEN lacks issue-write scope (403 on comment post). Gitea Actions
auto-provides a repo-scoped per-run token as secrets.GITEA_TOKEN — no
user-managed secret needed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PqGhF93UrpDNvXBLMjJENL
2026-07-03 13:39:59 +01:00
TudorandClaude Fable 5 d0895c71df fix(ci): AI review via Claude Code CLI; reuse REGISTRY_TOKEN for PR comments
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 9m39s
PR Checks / Backend Smoke (pull_request) Successful in 5s
PR Checks / Build Backend (no push) (pull_request) Successful in 16s
PR Checks / Build Frontend (no push) (pull_request) Successful in 42s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 9s
PR Checks / AI Code Review (Claude) (pull_request) Failing after 27s
- ai_review.py now pipes the diff through headless Claude Code (claude -p,
  --output-format json) authenticated with CLAUDE_CODE_OAUTH_TOKEN from
  'claude setup-token' — subscription auth, no Anthropic API billing
- stdlib-only script (urllib instead of requests/anthropic)
- PR comments posted with the existing REGISTRY_TOKEN secret; the separate
  GITEA_TOKEN secret is no longer needed

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PqGhF93UrpDNvXBLMjJENL
2026-07-03 08:28:20 +01:00
TudorandClaude Fable 5 4a52735356 feat(sdlc): staging environment + automated staging→prod pipeline
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 10m0s
PR Checks / Backend Smoke (pull_request) Successful in 48s
PR Checks / Build Backend (no push) (pull_request) Successful in 17s
PR Checks / Build Frontend (no push) (pull_request) Successful in 41s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 24s
- pr-checks.yml: PR gate — frontend typecheck+jest, backend import smoke,
  image builds (no push), Claude AI review posted as PR comment (severe
  findings block merge)
- deploy.yml (replaces build-and-push.yml): merge to main builds+pushes
  images tagged sha-<sha>/staging, deploys the staging Portainer stack via
  webhook, runs Playwright E2E journeys against staging, then retags the
  verified images :prod (previous kept as :prod-previous) and deploys prod
- docker-compose.portainer.staging.yml: second Portainer stack — :staging
  images, sc_staging_* names, own macvlan IPs, Airflow on 8081; data
  bootstrapped from source via the staging Airflow DAGs
- prod compose now pins :prod instead of :latest (only the promotion step
  moves it; :latest is no longer published)
- e2e/: 6 Playwright journeys (search, postcode, detail, compare, rankings)
  driven by BASE_URL — the promotion gate
- scripts/ci/ai_review.py: Claude review with structured JSON findings
- docs/DEPLOY.md: full SDLC doc incl. one-time setup checklist and rollback
- replaced removed 'next lint' with tsc typecheck; fixed stale jest tests
  (slug URLs, N/A formatting, stable trend, fake-timer setup)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PqGhF93UrpDNvXBLMjJENL
2026-07-03 06:50:40 +01:00