Ship-dark feature flags, with the last-distance-offered feature as the first thing behind one.
Merging this changes nothing visible in any environment. Every flag defaults off, and UNLEASH_URL is unset until someone sets it.
Why
Today a feature is either on main and live, or on a branch. That forces long-lived branches for anything unfinished and makes every promotion an all-or-nothing decision about everything queued behind it.
Self-hosted Unleash in its own Portainer stack — belonging to neither app stack, so a staging redeploy cannot disturb production's flags. One instance serves both: OSS Unleash ships development and production environments with environment-scoped client tokens, so a feature can be on in staging, where the E2E journeys exercise it, while production stays dark.
FastAPI holds the only SDK. Next reads /api/flags server-side, so package.json gains no Unleash dependency and no polling client competes with ISR.
Fail-closed everywhere. Unset config, unreachable server, a client that throws, an undeclared name — all False. Nothing here can 500 a request path or stop the API booting.
Unleash holds the state; backend/flags.py holds the list. The SDK evaluates an unknown flag to False, so without a registry that is an undeclared false, indistinguishable from a typo in a flag name.
Two findings worth reading
No webhook — the seven-day premise was wrong. The spec originally specified two Unleash webhooks and a revalidateTag purge. Next actually uses the lowestrevalidate among a route's fetches, not the segment value: school pages effectively revalidate every 5 minutes (fetchSchoolDetails at 300s), place pages every hour. Flags flip monthly, by hand. That deleted two webhook integrations, a revalidate route, a secret-in-query-string scheme, an idempotency requirement, and a rule that every fetch carry a cache tag — the part most likely to rot as fetches are added.
/api/flags would have been public.app/api/[...path]/route.ts forwards everything under /api/, so the endpoint would have published the name and state of every unreleased feature. Now denied on an exact first-segment match — not a prefix, so /api/flagship does not go down with /api/flags.
First consumer
admission_distance is on main, live on staging, and has never reached production (/api/schools/100010 there carries no such key). It needs exactly one gate, at the API: DistanceSection already returns null on a missing field, and only 57 LAs publish cut-offs, so the off-path is the commonest path on the site.
The field is absent, not null — /api/schools/ is public and unauthenticated, so a field left in the payload is a published field. That is the reasoning already recorded in c9a1892. The two are also different claims: null says this school has no cut-off; absent says cut-offs are not being published at all.
Verified CI's exact install path on Python 3.12 — pip install -r requirements.txt, import smoke test, suite green
New E2E journeys run against staging: the on-gate passes, the off-gate skips, the existing eight distance journeys are unaffected
The distance journeys previously skipped when no school had a figure, so a feature supposed to be on and silently broken would show as a green run full of skips. The new gate fails in that case.
Lifecycle
A test fails any flag older than 90 days — 2026-11-21 for this one. It fails on whatever PR is open at the time, which is the mechanism, not a bug in it. The fix is to delete the flag and the branch it guards.
Before this can be used
Task 1 of the plan ends with a manual step I cannot do: deploy docker-compose.portainer.unleash.yml in Portainer, create one client token per environment, and set UNLEASH_URL / UNLEASH_API_TOKEN in each app stack. Runbook is in docs/DEPLOY.md. Until then every flag is off, which is the correct dark state.
Known risk, named in the spec
Production gains a homelab dependency. If Unleash is unreachable when a backend cold-starts with an empty cache, every flag is False and any released feature disappears. The fcache volume covers restarts; the 90-day rule bounds the exposure. It is the reason flags must be retired rather than left on indefinitely.
Ship-dark feature flags, with the last-distance-offered feature as the first thing behind one.
**Merging this changes nothing visible in any environment.** Every flag defaults off, and `UNLEASH_URL` is unset until someone sets it.
## Why
Today a feature is either on `main` and live, or on a branch. That forces long-lived branches for anything unfinished and makes every promotion an all-or-nothing decision about everything queued behind it.
## Design
Spec: `docs/superpowers/specs/2026-08-23-feature-flags-design.md` · Plan: `docs/superpowers/plans/2026-08-23-feature-flags.md`
Self-hosted Unleash in **its own Portainer stack** — belonging to neither app stack, so a staging redeploy cannot disturb production's flags. One instance serves both: OSS Unleash ships `development` and `production` environments with environment-scoped client tokens, so a feature can be **on in staging, where the E2E journeys exercise it, while production stays dark**.
**FastAPI holds the only SDK.** Next reads `/api/flags` server-side, so `package.json` gains no Unleash dependency and no polling client competes with ISR.
**Fail-closed everywhere.** Unset config, unreachable server, a client that throws, an undeclared name — all `False`. Nothing here can 500 a request path or stop the API booting.
**Unleash holds the state; `backend/flags.py` holds the list.** The SDK evaluates an unknown flag to `False`, so without a registry that is an *undeclared* false, indistinguishable from a typo in a flag name.
## Two findings worth reading
**No webhook — the seven-day premise was wrong.** The spec originally specified two Unleash webhooks and a `revalidateTag` purge. Next actually uses the **lowest** `revalidate` among a route's fetches, not the segment value: school pages effectively revalidate every **5 minutes** (`fetchSchoolDetails` at 300s), place pages every **hour**. Flags flip monthly, by hand. That deleted two webhook integrations, a revalidate route, a secret-in-query-string scheme, an idempotency requirement, and a rule that every fetch carry a cache tag — the part most likely to rot as fetches are added.
**`/api/flags` would have been public.** `app/api/[...path]/route.ts` forwards everything under `/api/`, so the endpoint would have published the name and state of every unreleased feature. Now denied on an exact first-segment match — not a prefix, so `/api/flagship` does not go down with `/api/flags`.
## First consumer
`admission_distance` is on `main`, live on staging, and **has never reached production** (`/api/schools/100010` there carries no such key). It needs exactly one gate, at the API: `DistanceSection` already returns `null` on a missing field, and only 57 LAs publish cut-offs, so the off-path is the commonest path on the site.
The field is **absent, not null** — `/api/schools/` is public and unauthenticated, so a field left in the payload is a published field. That is the reasoning already recorded in `c9a1892`. The two are also different claims: null says *this school has no cut-off*; absent says *cut-offs are not being published at all*.
## Testing
- 133 backend, 269 frontend, `tsc` clean, `next build` green, 93 E2E collected (was 90)
- Verified CI's exact install path on Python 3.12 — `pip install -r requirements.txt`, import smoke test, suite green
- New E2E journeys run against staging: the on-gate passes, the off-gate skips, the existing eight distance journeys are unaffected
The distance journeys previously skipped when no school had a figure, so a feature *supposed* to be on and silently broken would show as a green run full of skips. The new gate fails in that case.
## Lifecycle
A test fails any flag older than 90 days — **2026-11-21** for this one. It fails on whatever PR is open at the time, which is the mechanism, not a bug in it. The fix is to delete the flag and the branch it guards.
## Before this can be used
Task 1 of the plan ends with a manual step I cannot do: deploy `docker-compose.portainer.unleash.yml` in Portainer, create one **client** token per environment, and set `UNLEASH_URL` / `UNLEASH_API_TOKEN` in each app stack. Runbook is in `docs/DEPLOY.md`. Until then every flag is off, which is the correct dark state.
## Known risk, named in the spec
Production gains a homelab dependency. If Unleash is unreachable when a backend cold-starts with an empty cache, every flag is `False` and any *released* feature disappears. The fcache volume covers restarts; the 90-day rule bounds the exposure. It is the reason flags must be retired rather than left on indefinitely.
Unleash self-hosted in its own Portainer stack, with FastAPI holding the
only SDK and Next reading flags through a tagged fetch.
The two hard parts are consequences of putting flag state in a service
rather than the repo: main stops being the whole truth about what is on,
and a flag can now change without the deploy that would have cleared the
caches. A code-declared registry bounds the first; webhook-driven
revalidateTag handles the second.
Cache tagging is deliberately coarse — every server fetch carries the
flags tag, not just the flags fetch itself. The first consumer proves
why: admission_distance changes the shape of /api/schools/{urn}, so a
narrow purge would leave ~25,000 school pages serving the pre-flip
render for a week, invisibly.
First consumer is the last-distance-offered feature, which is on main
and staging and has never reached production. It needs one gate, at the
API, because the frontend already no-ops on a missing field.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
Next uses the LOWEST revalidate among a route's fetches, not the segment
value. School pages fetch school details at 300s and place pages fetch
national averages at 3600s, so the effective ISR period is five minutes
and one hour respectively — not the seven days the segment declares.
A flag flip therefore propagates on its own, well inside the monthly,
by-hand cadence these flags are for. That deletes two webhook
integrations, a revalidate route, a secret-in-query-string scheme, an
idempotency requirement, and the rule that every fetch carry a cache
tag — which was the part most likely to rot as fetches are added.
Two constraints survive: a flag must never gate content on a
force-static page, because app/admissions never revalidates; and a
route-family flag must rebuild the sitemap, deferred with the route case
since no flag in scope touches it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
Task 1 is the Unleash stack and ends with a human step — the Portainer
deploy and the token generation cannot be automated from here. Nothing
else blocks on it: an unset UNLEASH_URL means every flag is False, which
is what local development and CI get, so the whole suite runs without a
flag server existing.
Self-review caught three defects in the plan itself. get_supplementary_data
takes (db, urn), not (urn), and the test DataFrame was minimised to the
point where the endpoint would have failed for reasons unrelated to
flags — both now copy the known-good shape from test_school_details.py.
The proxy test needs the node jest environment, since NextRequest wants
Fetch API globals jsdom does not provide. And the e2e off-state check
hardcoded a URN, so a 404 page would have satisfied it without proving
anything.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
Its own Portainer stack, belonging to neither application stack: a
staging redeploy must not be able to disturb production's flag state.
One instance serves both. OSS Unleash ships development and production
environments with environment-scoped client tokens, so the same flag
holds independent state in each — which is what lets a feature be on in
staging, where the E2E journeys exercise it, while production stays dark.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
Unleash holds flag state; it does not hold the list of flags. REGISTRY
is that list, because the SDK evaluates an unknown flag to False and
without a registry that is an undeclared False — indistinguishable from
a typo in a flag name.
Fail-closed throughout, and never raises: an unset UNLEASH_URL, an
unreachable server, a client that throws, an undeclared name — all
False. A flag layer that can 500 a request path or stop the API booting
is worse than one that is switched off.
Every flag defaults to False, with no per-flag override, because a flag
that defaults on is a kill switch and this is deliberately not one.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
The endpoint and its exposure control ship together on purpose. The
moment /api/flags exists, app/api/[...path] forwards it — and the
response names every unreleased feature the codebase knows about, along
with whether it is on. Publishing that is the opposite of shipping dark.
Denied on an exact first-segment match, not a prefix, so /api/flagship
does not go down with /api/flags. Next reads the endpoint server-side
over the Docker network, which never transits the public proxy.
jest.setup.js now guards its browser globals. It runs for every suite,
including the one that declares @jest-environment node to exercise the
route handler — NextRequest needs Fetch API globals jsdom lacks, and
there is no window there to define matchMedia on.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
Ships without a consumer, deliberately. The first flag needs none — the
backend withholds the field and the page follows — but 'UI elements on
existing pages' is one of the three surfaces this capability exists for,
and a flag layer that cannot gate one is incomplete.
Never throws: an unreadable flag is a dark one, which matches the
backend's fail-closed default. A page that 500s because the flags
endpoint blinked would be a worse outcome than a hidden feature.
Reading flags pins the calling route to a 300s ISR floor, since Next
takes the lowest revalidate among a route's fetches. That matches what
/school/[slug] already sits at, and it is the same property that makes a
flip propagate without a webhook.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
One gate, at the source. The frontend needs no change: DistanceSection
already returns null when distance_m is missing, and the admissions
block already conditions on (admissions || admissionDistance). Only 57
local authorities publish cut-offs, so the off-path is the commonest
path on the site and is well covered already.
Absent, not null. /api/schools/ is public and unauthenticated, so a
field left in the payload is a published field — the reasoning already
recorded in c9a1892 when history was withheld. The two are also
different claims: null says this school has no cut-off, absent says
cut-offs are not being published at all. The frontend type now says so.
The feature is on main and live on staging and has never reached
production, which is what makes it the right first consumer: the flag
lets the code promote without the feature appearing.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
The existing distance journeys all skip when no school has a published
figure, which is right when the feature is off — and wrong when it is
supposed to be on and is silently broken, because that shows up as a
green run full of skips. The new gate fails in exactly that case.
Feature state is read from the data, not from /api/flags: the public
proxy denies that path on purpose, since it names unreleased features.
Presence of the admission_distance key is the observable effect.
Verified against staging, where the feature is currently on: the on-gate
passes, the off-gate skips, the existing eight distance journeys are
unaffected.
One honest caveat — the /api/flags check passes on staging today because
that image predates the endpoint, not because the denylist works. The
denylist itself is covered by the jest unit test; this is defence in
depth and becomes a real assertion once deployed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
Both variables default to empty, so an environment without Unleash has
every flag off — the correct dark state rather than a boot failure.
The cache volume is the mitigation for the one real regression risk in
this design: the SDK evaluates everything False until it syncs, so a
backend cold-starting with an empty cache while Unleash is unreachable
would make a *released* feature disappear. On a named volume the disk
cache survives a restart.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
This PR adds a self-hosted Unleash feature-flag system: a new backend/flags.py registry with fail-closed evaluation, an internal-only /api/flags endpoint (protected by both network topology and a Next.js proxy denylist), a new Portainer stack for Unleash, and gates the existing admission_distance field behind the first flag. The implementation is careful and well-tested — the field is withheld at the source (absent key, not null), the proxy denylist matches exactly on the first path segment, network isolation already keeps the backend off the macvlan/public network, and CI installs the new UnleashClient dependency via requirements.txt. No severe issues found.
🟡 Minor
backend/flags.py: init() calls the synchronous UnleashClient.initialize_client() during FastAPI's lifespan startup with no explicit request_timeout override. If UNLEASH_URL is set but the Unleash server is slow or unreachable (network partition, wrong IP), this can block app startup for the SDK's default HTTP timeout on every container start/restart, slowing rollouts and health-check readiness even though the try/except prevents an outright crash.
backend/flags.py: The Unleash client started in init() is never stopped/destroyed on lifespan shutdown, leaving its background polling thread running past app shutdown in the same process (mostly harmless under Docker's process-per-container model, but sloppy for any in-process restart/hot-reload scenario).
## 🤖 AI Code Review (Claude Code)
This PR adds a self-hosted Unleash feature-flag system: a new backend/flags.py registry with fail-closed evaluation, an internal-only /api/flags endpoint (protected by both network topology and a Next.js proxy denylist), a new Portainer stack for Unleash, and gates the existing admission_distance field behind the first flag. The implementation is careful and well-tested — the field is withheld at the source (absent key, not null), the proxy denylist matches exactly on the first path segment, network isolation already keeps the backend off the macvlan/public network, and CI installs the new UnleashClient dependency via requirements.txt. No severe issues found.
### 🟡 Minor
- **backend/flags.py**: init() calls the synchronous UnleashClient.initialize_client() during FastAPI's lifespan startup with no explicit request_timeout override. If UNLEASH_URL is set but the Unleash server is slow or unreachable (network partition, wrong IP), this can block app startup for the SDK's default HTTP timeout on every container start/restart, slowing rollouts and health-check readiness even though the try/except prevents an outright crash.
- **backend/flags.py**: The Unleash client started in init() is never stopped/destroyed on lifespan shutdown, leaving its background polling thread running past app shutdown in the same process (mostly harmless under Docker's process-per-container model, but sloppy for any in-process restart/hot-reload scenario).
tudor
merged commit e953ee7c5f into main2026-08-23 11:34:56 +00:00
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Ship-dark feature flags, with the last-distance-offered feature as the first thing behind one.
Merging this changes nothing visible in any environment. Every flag defaults off, and
UNLEASH_URLis unset until someone sets it.Why
Today a feature is either on
mainand live, or on a branch. That forces long-lived branches for anything unfinished and makes every promotion an all-or-nothing decision about everything queued behind it.Design
Spec:
docs/superpowers/specs/2026-08-23-feature-flags-design.md· Plan:docs/superpowers/plans/2026-08-23-feature-flags.mdSelf-hosted Unleash in its own Portainer stack — belonging to neither app stack, so a staging redeploy cannot disturb production's flags. One instance serves both: OSS Unleash ships
developmentandproductionenvironments with environment-scoped client tokens, so a feature can be on in staging, where the E2E journeys exercise it, while production stays dark.FastAPI holds the only SDK. Next reads
/api/flagsserver-side, sopackage.jsongains no Unleash dependency and no polling client competes with ISR.Fail-closed everywhere. Unset config, unreachable server, a client that throws, an undeclared name — all
False. Nothing here can 500 a request path or stop the API booting.Unleash holds the state;
backend/flags.pyholds the list. The SDK evaluates an unknown flag toFalse, so without a registry that is an undeclared false, indistinguishable from a typo in a flag name.Two findings worth reading
No webhook — the seven-day premise was wrong. The spec originally specified two Unleash webhooks and a
revalidateTagpurge. Next actually uses the lowestrevalidateamong a route's fetches, not the segment value: school pages effectively revalidate every 5 minutes (fetchSchoolDetailsat 300s), place pages every hour. Flags flip monthly, by hand. That deleted two webhook integrations, a revalidate route, a secret-in-query-string scheme, an idempotency requirement, and a rule that every fetch carry a cache tag — the part most likely to rot as fetches are added./api/flagswould have been public.app/api/[...path]/route.tsforwards everything under/api/, so the endpoint would have published the name and state of every unreleased feature. Now denied on an exact first-segment match — not a prefix, so/api/flagshipdoes not go down with/api/flags.First consumer
admission_distanceis onmain, live on staging, and has never reached production (/api/schools/100010there carries no such key). It needs exactly one gate, at the API:DistanceSectionalready returnsnullon a missing field, and only 57 LAs publish cut-offs, so the off-path is the commonest path on the site.The field is absent, not null —
/api/schools/is public and unauthenticated, so a field left in the payload is a published field. That is the reasoning already recorded inc9a1892. The two are also different claims: null says this school has no cut-off; absent says cut-offs are not being published at all.Testing
tscclean,next buildgreen, 93 E2E collected (was 90)pip install -r requirements.txt, import smoke test, suite greenThe distance journeys previously skipped when no school had a figure, so a feature supposed to be on and silently broken would show as a green run full of skips. The new gate fails in that case.
Lifecycle
A test fails any flag older than 90 days — 2026-11-21 for this one. It fails on whatever PR is open at the time, which is the mechanism, not a bug in it. The fix is to delete the flag and the branch it guards.
Before this can be used
Task 1 of the plan ends with a manual step I cannot do: deploy
docker-compose.portainer.unleash.ymlin Portainer, create one client token per environment, and setUNLEASH_URL/UNLEASH_API_TOKENin each app stack. Runbook is indocs/DEPLOY.md. Until then every flag is off, which is the correct dark state.Known risk, named in the spec
Production gains a homelab dependency. If Unleash is unreachable when a backend cold-starts with an empty cache, every flag is
Falseand any released feature disappears. The fcache volume covers restarts; the 90-day rule bounds the exposure. It is the reason flags must be retired rather than left on indefinitely.Unleash self-hosted in its own Portainer stack, with FastAPI holding the only SDK and Next reading flags through a tagged fetch. The two hard parts are consequences of putting flag state in a service rather than the repo: main stops being the whole truth about what is on, and a flag can now change without the deploy that would have cleared the caches. A code-declared registry bounds the first; webhook-driven revalidateTag handles the second. Cache tagging is deliberately coarse — every server fetch carries the flags tag, not just the flags fetch itself. The first consumer proves why: admission_distance changes the shape of /api/schools/{urn}, so a narrow purge would leave ~25,000 school pages serving the pre-flip render for a week, invisibly. First consumer is the last-distance-offered feature, which is on main and staging and has never reached production. It needs one gate, at the API, because the frontend already no-ops on a missing field. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj🤖 AI Code Review (Claude Code)
This PR adds a self-hosted Unleash feature-flag system: a new backend/flags.py registry with fail-closed evaluation, an internal-only /api/flags endpoint (protected by both network topology and a Next.js proxy denylist), a new Portainer stack for Unleash, and gates the existing admission_distance field behind the first flag. The implementation is careful and well-tested — the field is withheld at the source (absent key, not null), the proxy denylist matches exactly on the first path segment, network isolation already keeps the backend off the macvlan/public network, and CI installs the new UnleashClient dependency via requirements.txt. No severe issues found.
🟡 Minor