feat(suggest): school autosuggest, and the rate-limit fix it needed first #127

Merged
tudor merged 11 commits from feat/school-autosuggest into main 2026-08-26 19:58:58 +00:00
Showing only changes of commit e651dd0d65 - Show all commits
@@ -70,20 +70,30 @@ single-process uvicorn backend that filters a 25,000-row DataFrame in-process.
Correct per-user keying removes that throttle: the origin becomes reachable at
60/min *per user* rather than 60/min in total.
So the keying fix ships **with a global ceiling**, not instead of one:
slowapi's application-wide `default_limits`, set above any plausible real load
but below what would flood a single worker. **3000/minute**, against a
per-user suggest limit of 120/minute — so twenty-five simultaneously active
searchers are unaffected, while a runaway client or a scraper still meets a
wall.
Per-user fairness and origin protection are different jobs. Conflating them
is what produced the current behaviour, and the fix must not quietly do it
again in the other direction.
Per-user fairness and origin protection are different jobs and need different
limits. Conflating them is what produced the current behaviour.
**The global ceiling does not go in this app.** An earlier draft of this
section specified one via slowapi's `default_limits`. Reading the library
shows that would not have worked, twice over: `default_limits` and
`application_limits` are both evaluated with the same `key_func`, so they are
per-client across all routes rather than global; and `application_limits` are
only applied `if in_middleware`, while this app installs no `SlowAPIMiddleware`
at all. Expressing a genuine global cap would take a second `Limiter` with a
constant key plus that middleware — two mechanisms to keep correct, for a
protection this layer is the wrong place for.
Both numbers are estimates, not measurements. They are a starting point to
revisit against real traffic once the keying is correct enough for real
traffic to be visible — which it is not today, because everyone shares one
bucket.
Cloudflare is already in the request path on both environments and does
edge-level rate limiting properly, before traffic reaches a single-process
origin at all. That is where a global ceiling belongs, and it is a dashboard
change rather than code. Flagged as a follow-up, deliberately not built here.
What ships instead is conservative per-user limits: the existing 60/minute
default is unchanged, and `/api/suggest` gets 120/minute. Both are estimates
rather than measurements, and they are a starting point to revisit once the
keying is correct enough for real per-user traffic to be visible — which it
is not today, because everyone shares one bucket.
## 2. `GET /api/suggest`
@@ -115,7 +125,7 @@ GET /api/suggest?q=<query>&limit=8
- **Rate limit `120/minute`** per client, not the default 60. A 200 ms
debounce tops out near 5 requests/second while someone is actively typing,
but averages far below that across a real search; 120 leaves headroom for
bursts without letting one client alone consume the global ceiling.
bursts while still bounding one client.
- **Local authority is part of the payload, not decoration.** There are many
schools called "St Mary's"; a suggestion list without the authority is
unusable for exactly the queries autosuggest is meant to serve.
@@ -210,9 +220,13 @@ effect, since `/api/flags` is denied to the public.
## 7. Risks
**Removing the accidental throttle.** Covered in §1. The global ceiling is the
mitigation, and 1200/minute is a guess that should be revisited against real
traffic rather than treated as a considered number.
**Removing the accidental throttle.** Covered in §1. Correct per-user keying
means the origin is reachable at 60/minute *per user* where it was 60/minute
in total, and no in-app global cap replaces it — that job goes to Cloudflare,
which is not done as part of this change. Until it is, a determined caller
with many source addresses can put more load on a single-process origin than
they can today. Against this site's traffic that is a theoretical risk rather
than a live one, but it is a real one and it is the price of the fix.
**Cloudflare bypass.** If the origin is reachable without passing through
Cloudflare, `CF-Connecting-IP` is absent and the `X-Forwarded-For` fallback is