Closes the reliability gaps that docs/ARCHITECTURE.md had been carrying as a "Known limitations" paragraph, and replaces that paragraph with a description of what now happens.
The problems
Promotion could ship an untested image. Staging health polling asked only whether something answered 200 at the base URL. It could not distinguish the new deployment from the old, so journeys could pass against the previous release. Overlapping merges could move the staging tags underneath a run in flight.
Reload emptied the API on failure./api/admin/reload cleared the caches before rebuilding, so any failure left the backend serving nothing, and requests arriving mid-reload saw a half-swapped state.
The search index published unvalidated.sync_typesense.py never read an import response before swapping the alias; a partial import went live. Two overlapping DAG runs could prune each other's collections.
Search silently truncated. Typesense returned at most one page of hits, and used [] for both "no matches" and "search is down" — so a genuine empty result fell back to substring matching.
Outages looked like absence. The home page rendered its empty state on any fetch failure; school pages called notFound(), telling visitors and crawlers that a real school had ceased to exist.
Stale results appended themselves. "Load more" and the map fetch resolved against whatever state existed when they returned, so results from an abandoned search merged into the new ones.
The changes
Each staging run mints a build ID and stamps all three images with the commit and that ID — as labels, and for frontend/backend as a build-time JSON file that environment overrides cannot rewrite. /release.json reports both identities uncached. scripts/ci/release.py polls for the expected pair before and after the journeys, then tags the captured build digests verified-<sha>. Promotion resolves those verified tags to immutable digests, revalidates their labels, and refuses a mixed or incomplete set before any :prod tag moves. The staging workflow shares one concurrency group with cancellation disabled.
Backend reload now builds and validates the replacement frames, place registry, reverse index and sitemaps off the request loop, then publishes them in one synchronous step under a lock; failure returns 503 and keeps the previous data. Index publication checks every import response and the final document count before upserting the alias, under a PostgreSQL advisory lock, keeping the previous collection as a rollback pointer. Frontend failures reach a retryable error boundary, and every client fetch carries an AbortController plus a search-scope check.
New suites: dataset publication, search completeness, index publication, release script, workflow shape, stale-fetch guards, data-failure surfacing.
New Playwright journeys check deployed release identity and stale pagination.
PR checks now run pipeline/tests and scripts/ci/tests alongside the backend suite.
Not proven locally, and deliberately so: registry credentials, Portainer behaviour, proxy routing and any deployed image. Those need the staging run. Per docs/DEPLOY.md, the concurrency key needs Gitea ≥ 1.26 (the server reported 1.27.3), and the first rollout matters: older green commits have no verified tags or build identities, so a commit built by this workflow must be the first thing promoted through the gate. Production promotion remains a separate human action.
Closes the reliability gaps that `docs/ARCHITECTURE.md` had been carrying as a "Known limitations" paragraph, and replaces that paragraph with a description of what now happens.
## The problems
- **Promotion could ship an untested image.** Staging health polling asked only whether *something* answered 200 at the base URL. It could not distinguish the new deployment from the old, so journeys could pass against the previous release. Overlapping merges could move the staging tags underneath a run in flight.
- **Reload emptied the API on failure.** `/api/admin/reload` cleared the caches before rebuilding, so any failure left the backend serving nothing, and requests arriving mid-reload saw a half-swapped state.
- **The search index published unvalidated.** `sync_typesense.py` never read an import response before swapping the alias; a partial import went live. Two overlapping DAG runs could prune each other's collections.
- **Search silently truncated.** Typesense returned at most one page of hits, and used `[]` for both "no matches" and "search is down" — so a genuine empty result fell back to substring matching.
- **Outages looked like absence.** The home page rendered its empty state on any fetch failure; school pages called `notFound()`, telling visitors and crawlers that a real school had ceased to exist.
- **Stale results appended themselves.** "Load more" and the map fetch resolved against whatever state existed when they returned, so results from an abandoned search merged into the new ones.
## The changes
Each staging run mints a build ID and stamps all three images with the commit and that ID — as labels, and for frontend/backend as a build-time JSON file that environment overrides cannot rewrite. `/release.json` reports both identities uncached. `scripts/ci/release.py` polls for the expected pair before *and* after the journeys, then tags the captured build digests `verified-<sha>`. Promotion resolves those verified tags to immutable digests, revalidates their labels, and refuses a mixed or incomplete set before any `:prod` tag moves. The staging workflow shares one concurrency group with cancellation disabled.
Backend reload now builds and validates the replacement frames, place registry, reverse index and sitemaps off the request loop, then publishes them in one synchronous step under a lock; failure returns 503 and keeps the previous data. Index publication checks every import response and the final document count before upserting the alias, under a PostgreSQL advisory lock, keeping the previous collection as a rollback pointer. Frontend failures reach a retryable error boundary, and every client fetch carries an `AbortController` plus a search-scope check.
## Verification
- 219 backend/pipeline/CI tests pass; 439 frontend tests pass; `tsc --noEmit` clean.
- New suites: dataset publication, search completeness, index publication, release script, workflow shape, stale-fetch guards, data-failure surfacing.
- New Playwright journeys check deployed release identity and stale pagination.
- PR checks now run `pipeline/tests` and `scripts/ci/tests` alongside the backend suite.
**Not proven locally, and deliberately so:** registry credentials, Portainer behaviour, proxy routing and any deployed image. Those need the staging run. Per `docs/DEPLOY.md`, the concurrency key needs Gitea ≥ 1.26 (the server reported 1.27.3), and the **first rollout** matters: older green commits have no verified tags or build identities, so a commit built *by this workflow* must be the first thing promoted through the gate. Production promotion remains a separate human action.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
The old sync created a collection, imported batches without reading a
single import response, and swapped the alias regardless. A partial
import published a half-empty index, and two overlapping DAG runs could
prune each other's collections.
Publication now checks every import response and the final document
count before upserting the alias, and holds a session-scoped advisory
lock across the read and the publish so concurrent runs serialise.
Cleanup keeps the previous collection as a rollback pointer and is
best-effort: an uncertain alias response must never delete what might
still be live.
Also parses the Typesense URL properly instead of splitting on colons,
which mangled any host carrying a scheme and a default port.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Reload cleared the caches first and rebuilt afterwards, so any failure
left the API serving nothing, and requests arriving mid-reload saw a
half-swapped state. It now builds and validates the replacement frames,
place registry, reverse index and sitemaps off the request loop, then
publishes them in one synchronous step under a lock. A failed reload
returns 503 and keeps the previous data. Sitemap regeneration takes the
same path rather than clearing the live registry up front.
Typesense search returned at most one page of hits and used an empty
list for both "no matches" and "search is down", so a genuine empty
result silently fell back to substring matching. It now pages through
every candidate and returns None only on failure; the fallback matches
literally, since a query containing regex metacharacters used to throw.
Empty datasets answer 503 rather than 200-with-nothing or a misleading
404, so callers can tell an outage from an absent school. Adds
/api/release, which reports the build identity baked into the image.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The home page caught every fetch failure and rendered its empty state,
so a backend outage looked like a site with no schools in it. School
pages turned any error into notFound(), which told visitors — and
crawlers — that a real school had ceased to exist. Place fetches did the
same by returning [] and null. Failures now reach a retryable error
boundary; only a genuine 404 still calls notFound().
"Load more" and the map fetch resolved against whatever state existed
when they returned, so results from an abandoned search appended
themselves to the new ones. Each fetch now carries an AbortController
and checks that its search scope is still current before touching state.
The map only records its cache key on success, so a failed load retries
instead of pinning the stale marker set.
Jest ignored .next/, whose build output otherwise shadowed real suites.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Staging health polling asked only whether something answered HTTP 200 at
the base URL. It could not tell the new deployment from the old one, so
journeys could pass against the previous release, and concurrent merges
could move the staging tags underneath a run in flight.
Each staging run now mints a build ID and stamps all three images with
the commit and that ID, as labels and — for frontend and backend — as a
build-time JSON file that environment overrides cannot rewrite.
/release.json reports both identities uncached, and scripts/ci/release.py
polls for the expected pair before and after the journeys. Only then are
the captured build digests tagged verified-<sha>.
Promotion resolves those verified tags to immutable digests, revalidates
their labels, and refuses a mixed or incomplete set before any :prod tag
moves. The whole staging workflow shares one concurrency group with
cancellation disabled, so releases serialise.
The scripts are stdlib-only and unit-tested against mocked registry and
HTTP behaviour; PR checks now run the pipeline and CI suites too. The
runbook records what this cannot prove locally, and that the first
rollout needs a commit built by this workflow.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This PR adds an end-to-end release-identity gate (random build IDs baked into image labels/build-info.json, verified via a new scripts/ci/release.py, serialized staging deploys via a concurrency group, and digest-pinned production promotion), rewrites backend data publication to be atomic/validated (reload and sitemap regen now build everything off-thread and swap globals together, never leaving a half-updated or emptied state), fixes a real Typesense-vs-substring-fallback bug (empty result set was previously conflated with 'search unavailable'), makes Typesense search results complete via full pagination, and hardens the Next.js frontend against stale/aborted fetches and mis-classified errors (404 vs. real outages). The change is large but internally consistent, extensively covered by new unit/integration tests, and the one issue found is a minor performance concern rather than a correctness or safety defect.
🟡 Minor
backend/data_loader.py: search_schools_typesense now fetches every matching page from Typesense with no cap (previously bounded at 250). A common search term can match thousands of schools, causing dozens of sequential Typesense round trips per request, increasing latency/load under normal or adversarial search-as-you-type usage.
## 🤖 AI Code Review (Claude Code)
This PR adds an end-to-end release-identity gate (random build IDs baked into image labels/build-info.json, verified via a new scripts/ci/release.py, serialized staging deploys via a concurrency group, and digest-pinned production promotion), rewrites backend data publication to be atomic/validated (reload and sitemap regen now build everything off-thread and swap globals together, never leaving a half-updated or emptied state), fixes a real Typesense-vs-substring-fallback bug (empty result set was previously conflated with 'search unavailable'), makes Typesense search results complete via full pagination, and hardens the Next.js frontend against stale/aborted fetches and mis-classified errors (404 vs. real outages). The change is large but internally consistent, extensively covered by new unit/integration tests, and the one issue found is a minor performance concern rather than a correctness or safety defect.
### 🟡 Minor
- **backend/data_loader.py**: search_schools_typesense now fetches every matching page from Typesense with no cap (previously bounded at 250). A common search term can match thousands of schools, causing dozens of sequential Typesense round trips per request, increasing latency/load under normal or adversarial search-as-you-type usage.
Fetching every match kept scoped searches correct but left the number of
round trips in the caller's hands: a one-letter query, or a deliberately
broad one, walked the whole collection a page at a time.
Cap the candidate set at 1,000 URNs — four pages — and return the
relevance-ordered prefix when the ceiling is hit. That is still far more
than one page, so the API's own authority, phase and postcode filters
keep the matches they need, while latency and upstream load stay bounded.
A capped query is logged so a genuinely truncated search is visible.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
tudor
merged commit 7ab084dd3a into main2026-09-15 11:13:08 +00:00
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Closes the reliability gaps that
docs/ARCHITECTURE.mdhad been carrying as a "Known limitations" paragraph, and replaces that paragraph with a description of what now happens.The problems
/api/admin/reloadcleared the caches before rebuilding, so any failure left the backend serving nothing, and requests arriving mid-reload saw a half-swapped state.sync_typesense.pynever read an import response before swapping the alias; a partial import went live. Two overlapping DAG runs could prune each other's collections.[]for both "no matches" and "search is down" — so a genuine empty result fell back to substring matching.notFound(), telling visitors and crawlers that a real school had ceased to exist.The changes
Each staging run mints a build ID and stamps all three images with the commit and that ID — as labels, and for frontend/backend as a build-time JSON file that environment overrides cannot rewrite.
/release.jsonreports both identities uncached.scripts/ci/release.pypolls for the expected pair before and after the journeys, then tags the captured build digestsverified-<sha>. Promotion resolves those verified tags to immutable digests, revalidates their labels, and refuses a mixed or incomplete set before any:prodtag moves. The staging workflow shares one concurrency group with cancellation disabled.Backend reload now builds and validates the replacement frames, place registry, reverse index and sitemaps off the request loop, then publishes them in one synchronous step under a lock; failure returns 503 and keeps the previous data. Index publication checks every import response and the final document count before upserting the alias, under a PostgreSQL advisory lock, keeping the previous collection as a rollback pointer. Frontend failures reach a retryable error boundary, and every client fetch carries an
AbortControllerplus a search-scope check.Verification
tsc --noEmitclean.pipeline/testsandscripts/ci/testsalongside the backend suite.Not proven locally, and deliberately so: registry credentials, Portainer behaviour, proxy routing and any deployed image. Those need the staging run. Per
docs/DEPLOY.md, the concurrency key needs Gitea ≥ 1.26 (the server reported 1.27.3), and the first rollout matters: older green commits have no verified tags or build identities, so a commit built by this workflow must be the first thing promoted through the gate. Production promotion remains a separate human action.🤖 Generated with Claude Code
🤖 AI Code Review (Claude Code)
This PR adds an end-to-end release-identity gate (random build IDs baked into image labels/build-info.json, verified via a new scripts/ci/release.py, serialized staging deploys via a concurrency group, and digest-pinned production promotion), rewrites backend data publication to be atomic/validated (reload and sitemap regen now build everything off-thread and swap globals together, never leaving a half-updated or emptied state), fixes a real Typesense-vs-substring-fallback bug (empty result set was previously conflated with 'search unavailable'), makes Typesense search results complete via full pagination, and hardens the Next.js frontend against stale/aborted fetches and mis-classified errors (404 vs. real outages). The change is large but internally consistent, extensively covered by new unit/integration tests, and the one issue found is a minor performance concern rather than a correctness or safety defect.
🟡 Minor