# Feature Flags — Design **Date:** 2026-08-23 **Status:** approved for planning **First consumer:** the last-distance-offered feature (`admission_distance`) ## Goal Let work merge to `main` and deploy to production without becoming visible, so that releasing a feature stops being the same event as deploying it. The site has no way to do this today. A feature is either on `main` and live, or it is on a branch. That forces long-lived branches for anything not ready, and it makes every promotion to production an all-or-nothing decision about everything queued behind it. This is a **ship-dark** capability, not a kill switch. Flags are expected to flip on the order of once a month, by a person, deliberately. Nothing here is designed for flipping something off in seconds under pressure, and nothing here does percentage rollouts, user targeting or A/B tests — the site has no user identity to target. ## Decision: Unleash Flag state is held in a self-hosted [Unleash](https://www.getunleash.io/) instance (Apache-2.0), not in the repository. A lighter option was considered and rejected by the project owner: a typed registry in each runtime with environment-variable overrides set in the Portainer stack files, which would have needed no new container and kept flag state in git. The argument for Unleash is that it provides a UI and an audit log without a deploy, and that flags are expected to become an ongoing operational tool rather than an occasional one. Two consequences follow from choosing a service, and this design exists mostly to handle them: 1. **Flag state lives outside the repository.** `main` is no longer the whole truth about what is switched on. The registry in §2 exists to bound that. 2. **A flag can change without a deploy**, so nothing else clears the caches that a deploy would have cleared. §4 establishes how long a flip takes to become visible, and why that is short enough to need no extra mechanism. Also considered: Flagsmith (heavier — Django, Postgres and Redis), GrowthBook (requires MongoDB), and Flipt v2 (the closest conceptual fit, git-native, but now under the Fair Core Licence — source-available, not OSI open source). ## 1. Topology A third Portainer stack, `docker-compose.portainer.unleash.yml`, holding `unleashorg/unleash-server` and its own PostgreSQL 16. It is on the macvlan so both application stacks can reach it, and it belongs to neither of them — a staging redeploy must not be able to disturb production's flag state, and vice versa. One instance serves both environments. Open-source Unleash ships with `development` and `production` environments and environment-scoped client tokens, so the same flag holds independent state in each: staging's FastAPI carries a `development` token, production's carries a `production` one. That property is what makes ship-dark testable. A feature can be **on in staging and off in production** for as long as it takes, which means the `e2e/` journeys exercise it against staging while production stays unchanged. ## 2. The registry Unleash supplies flag *state* and the toggle UI. It does not supply the list of flags. `backend/flags.py` declares every flag the code knows about: ```python @dataclass(frozen=True) class Flag: name: str # identical in the registry, in Unleash, and in JSON description: str # one line: what turning this on reveals added: date # for the staleness test in §8 ``` **Every flag defaults to `False`.** There is no per-flag default field, because a flag that defaults on is not a ship-dark flag — it is a kill switch, and this design does not offer one. A single unconditional default also means the fallback path has no branching to get wrong. Three reasons the registry is not optional: - The Unleash SDK evaluates an unknown flag to `False`. Without a registry that is an *undeclared* false — indistinguishable from a typo in a flag name. - `/api/flags` needs a key set to return when Unleash is unreachable. It cannot enumerate flags it has never heard of. - A flag present in the Unleash UI but absent from the registry is orphaned, and should be visibly so rather than quietly authoritative. **Naming.** One string, used unchanged as the registry key, the Unleash flag name, and the JSON key in `/api/flags`. It is snake_case, matching the API's existing convention (`admission_distance`, `rwm_expected_pct`) and the mirrored types in `nextjs-app/lib/types.ts`. No case transformation anywhere, so there is no mapping layer to get wrong. ## 3. Read paths ### Backend `backend/flags.py` wraps `UnleashClient` behind `is_enabled(name: str) -> bool`. Fail-closed is the default rather than something added: the Python SDK evaluates every flag to `False` until it has synchronised with the server. An unfinished feature therefore stays hidden when Unleash is unreachable, which is the correct direction for ship-dark. The SDK's fcache directory is mounted on a named volume so a container restart during an Unleash outage keeps last-known state rather than reverting a released feature to dark. The registry default remains `False`, so the worst case is a feature disappearing, never one appearing. ### Frontend `nextjs-app/lib/flags.ts` exposes `getFlags(): Promise`, a single server-side fetch of `/api/flags` returning a typed record. Server components only — no flag value reaches the browser bundle, and `package.json` gains no Unleash dependency. The Unleash client library stays entirely inside the service that already owns every other piece of data the frontend renders. The cost, named plainly: a purely front-end flag must still be declared in a Python file. It is a flat data edit rather than programming, and the return is one list, so nobody has to ask which service knows about a given flag. ### `/api/flags` must not be publicly reachable `nextjs-app/app/api/[...path]/route.ts` proxies **everything** under `/api/` to FastAPI. Left alone, `https://www.schoolcompare.co.uk/api/flags` would return `{"admission_distance": false, ...}` — publishing the name and state of every unreleased feature, which defeats the purpose of shipping dark. The proxy therefore gains a denylist, and `flags` is on it: a request for a denied path returns 404 rather than being forwarded. Next's own `getFlags()` is unaffected because it calls `FASTAPI_URL` directly across the Docker network and never transits the public proxy. This is a general hole rather than a flags-specific one — the proxy will forward any future internal endpoint too — so the denylist is written as a named constant with a comment saying what belongs on it. ## 4. Propagation **Time-based revalidation is sufficient. There is no webhook.** An earlier draft of this section specified two Unleash webhooks and a `revalidateTag('flags')` purge, on the premise that pages cache for seven days. That premise was wrong, and checking it removed the most complex part of the design. Next uses the **lowest** `revalidate` among a route's fetches to set the revalidation frequency of the whole route — the segment-level `export const revalidate` does not override a lower value inside it. Measured against this codebase: | Page family | Segment | Lowest fetch | Effective | |---|---|---|---| | `/school/[slug]` | 604800 | `fetchSchoolDetails` at 300 | **5 minutes** | | `/schools/*` | 604800 | `fetchNationalAverages` at 3600 | **1 hour** | The Unleash SDK polls every 15 seconds, so a flip reaches school pages within about five minutes and place pages within the hour, unaided. Flags flip monthly, by hand, deliberately. That is fast enough. What this removes: two webhook integrations, a `/api/revalidate-flags` route, a shared-secret-in-a-query-string scheme, an idempotency requirement against duplicate and out-of-order delivery, and a rule that every fetch in `nextjs-app/lib/` carry a cache tag. None of it has to be built, maintained, or kept correct as new fetches are added. **If instant flips are ever wanted**, the webhook is the way to add them, and it is purely additive — nothing in this design has to change first. ### Two constraints this leaves behind **Never flag content on a `force-static` page.** `app/admissions/page.tsx` declares `export const dynamic = 'force-static'`, so it is baked at build time and never revalidates. A flag gating anything on such a page would not take effect until the next deploy, silently. If a flag ever needs to reach one, that page must first move to ISR. **A route-family flag still needs the sitemap rebuilt.** The sitemap is held in memory and rebuilt only at startup or via `POST /api/admin/regenerate-sitemap`. No flag in scope touches the sitemap (§6), so this is deferred with the route case rather than solved now — but a route flag must not ship without it, or the sitemap will advertise URLs that `notFound()`. ## 5. What "off" means, per surface | Surface | Off | |---|---| | Route | `notFound()`, **and** absent from the sitemap, **and** absent from nav | | UI element | Not rendered; surrounding page byte-identical to today | | API field | Key **absent**, not `null` | | API endpoint | 404, not 403 | The three parts of the route rule move together or not at all. Submitting URLs to Google that return 404 is the bug fixed in PR #124, and a flag is a new way to reintroduce it. An API field is withheld **at the source**, never rendered-but-hidden. The precedent is already set in this codebase by commit `c9a1892`: `/api/schools/` is public and unauthenticated, so leaving a withheld field in the payload hands the record to anyone who opens the network tab. ## 6. First consumer: `admission_distance` The last-distance-offered feature is merged to `main` and live on staging. Production has never received it: `/api/schools/100010` on production carries no `admission_distance` key, and no Distance section renders. It needs **exactly one gate** — `backend/app.py:809`, where the field is attached to the school payload: ```python "admission_distance": ( supplementary.get("admission_distance") if flags.is_enabled("admission_distance") else None ), ``` The frontend follows with no change. `DistanceSection` already returns `null` when `admission_distance?.distance_m == null`, and `PrimarySchoolSections` already conditions the admissions block on `(admissions || admissionDistance)`. The off-state is the commonest state on the site — only 57 local authorities publish cut-off distances at all — so it is well covered by construction. The flag does not touch the sitemap: school pages exist either way. Intended lifecycle: default off, so production receives the code dark on the next promotion; on in the `development` environment so staging keeps testing it; flipped on in `production` when the owner chooses. **This flag exercises two of the three surfaces** in §5 — API field and UI element. No route case ships with it. The route rule is specified but unproven until a route-shaped flag exists, and should be treated as such. ## 7. Testing **Backend unit.** The registry is well-formed; an unknown flag evaluates `False`; `/api/flags` returns every declared flag with its default when the SDK is unreachable; `admission_distance` is absent from the school payload when the flag is off and present when on. **Frontend unit.** `getFlags()` returns declared defaults when `/api/flags` fails, rather than throwing and taking the page with it. **E2E.** Journeys read `/api/flags` and gate flag-dependent assertions on it, matching the `test.skip` shape the suite already uses. One trap to avoid, worth stating because the existing distance journeys walk straight into it: they already skip when no school has a published figure, so with the flag off they would skip silently and the suite would go green. The gate must be explicit — **if `/api/flags` reports `admission_distance` on, then a school with a cut-off must be found**, converting a silent skip into a real assertion. ## 8. Lifecycle A flag is temporary scaffolding, and the failure mode of every flag system is accumulation. The registry records the date each flag was added, and a backend test fails any flag older than **90 days**. Removing a flag means deleting the registry entry, the branches that read it, and the flag in the Unleash UI. Unleash SDK usage metrics stay enabled, so the UI shows which flags are still being evaluated — the evidence needed to retire one safely. ## 9. Risks **Production gains a homelab dependency.** If Unleash is unreachable when a production container cold-starts with an empty cache, every flag evaluates `False` and any feature currently switched on disappears. The fcache volume covers restarts; the 90-day lifecycle rule bounds how long any feature is exposed to this. It is a real regression risk and the reason flags must be retired rather than left on indefinitely. **Flag state is not in git.** `main` no longer tells you what production is showing. The registry lists what *can* be flagged; only the Unleash UI says what *is*. This is inherent to the choice of a service. **A large promotion backlog exists.** Production is running the pre-SEO-programme build — no place pages, and a sitemap still declaring the apex host. The first promotion after this work ships that entire backlog. The flag isolates the distance feature from it and nothing else. ## Out of scope - Percentage rollouts, user targeting, A/B testing, and Unleash strategies beyond simple on/off. Flags are booleans. - Pipeline and dbt flags. Airflow and dbt are not flag consumers. - Client-side flag evaluation. Flags are server-side only. - Automatic flag removal. The staleness test reports; a person deletes.