fix(api): bound what a forged CF-Connecting-IP can buy
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 18s
PR Checks / Build Frontend (no push) (pull_request) Successful in 45s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 10s
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 18s
PR Checks / Build Frontend (no push) (pull_request) Successful in 45s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 10s
Code review, both findings valid. The design doc claimed Cloudflare "replaces the header, so a browser cannot forge it", and that only the X-Forwarded-For fallback was forgeable. That is true only for traffic that actually passed through Cloudflare, and nothing in this process can verify that it did. Reaching the origin directly, both headers are equally attacker-controlled — and rotating CF-Connecting-IP mints a fresh rate-limit bucket per request, defeating per-client limits on every endpoint including the DataFrame-heavy /api/schools. Against abuse that is worse than the shared bucket it replaced, which at least capped everyone together. So the ceiling comes back. I dropped it earlier arguing it belonged at Cloudflare; that argument assumed the keying was sound, and it is not. GlobalRateLimitMiddleware counts all /api/ traffic in a fixed window against a total, independent of client identity, outermost so it refuses before any work happens. Written by hand because slowapi cannot express a global cap: default_limits and application_limits are both keyed by key_func, and the latter needs middleware this app does not install. It does not make the header trustworthy — it makes trusting it survivable. The real fix is Authenticated Origin Pulls or an origin firewall, now documented in DEPLOY.md as the open gap it is. 127.0.0.1 is exempt: the healthcheck curls localhost from inside the container, and starving it would restart the container and turn a load spike into an outage loop. Keyed on the peer address, never the Host header, which the caller sets. Second finding: suggest_schools_typesense promised "never raises" while the parsing loop sat outside the try, so int(None) on a malformed document would have made a keystroke a 500. The loop now skips bad rows rather than dropping the whole list — and a hit with no document no longer becomes a suggestion pointing at /school/0. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
This commit is contained in:
1 parent
d2115364ae
commit
0fa1a292c7
8 files changed
+332
-50
No files matched your search
+76
-6
@@ -6,6 +6,7 @@ Uses real data from UK Government Compare School Performance downloads.
|
||||
|
||||
import hashlib
|
||||
import re
|
||||
import time
|
||||
from contextlib import asynccontextmanager
|
||||
from datetime import datetime, timezone
|
||||
from typing import Optional
|
||||
@@ -15,7 +16,7 @@ import pandas as pd
|
||||
from fastapi import FastAPI, HTTPException, Query, Request, Depends, Header
|
||||
from fastapi.middleware.cors import CORSMiddleware
|
||||
from fastapi.middleware.gzip import GZipMiddleware
|
||||
from fastapi.responses import FileResponse, Response
|
||||
from fastapi.responses import FileResponse, JSONResponse, Response
|
||||
from fastapi.staticfiles import StaticFiles
|
||||
from slowapi import Limiter, _rate_limit_exceeded_handler
|
||||
from slowapi.util import get_remote_address
|
||||
@@ -321,14 +322,79 @@ def client_key(request: Request) -> str:
|
||||
return get_remote_address(request)
|
||||
|
||||
|
||||
# Rate limiter. No in-app global ceiling: slowapi's default_limits and
|
||||
# application_limits are both keyed by key_func (so per-client, not global)
|
||||
# and the latter only applies with SlowAPIMiddleware installed, which this app
|
||||
# does not use. A global cap belongs at Cloudflare, which is already in the
|
||||
# path. See the spec's §1 for why that is deliberate.
|
||||
# Per-client limiter. Paired with the global ceiling below — the two do
|
||||
# different jobs and neither substitutes for the other.
|
||||
limiter = Limiter(key_func=client_key)
|
||||
|
||||
|
||||
# --- The ceiling no header can raise ----------------------------------------
|
||||
#
|
||||
# client_key trusts CF-Connecting-IP, and nothing in this process can tell an
|
||||
# edge-set header from an attacker-set one. That distinction can only be made
|
||||
# at Cloudflare, with Authenticated Origin Pulls or an origin firewall. A
|
||||
# caller reaching the origin directly could otherwise mint a fresh rate-limit
|
||||
# bucket per request and evade per-client limits entirely — which would make
|
||||
# correct keying a net regression against abuse, since the single shared bucket
|
||||
# it replaced at least capped everyone at 60/minute together.
|
||||
#
|
||||
# So per-client limits give fairness, and this gives the origin a hard total.
|
||||
# It does not make the header trustworthy; it bounds what trusting it can cost.
|
||||
# The header problem itself is closed at Cloudflare, not here.
|
||||
#
|
||||
# [window_start_monotonic, count], or None before the first request. A fixed
|
||||
# window is crude, which is right for a backstop: it has to be obviously
|
||||
# correct rather than fair.
|
||||
_global_window: Optional[list] = None
|
||||
|
||||
# The container healthcheck runs `curl http://localhost:80/api/data-info` from
|
||||
# inside the container. Starving it would fail the check, restart the
|
||||
# container, and turn a load spike into an outage loop — the ceiling exists to
|
||||
# protect the origin, not to kill it.
|
||||
_LOCAL_HOSTS = frozenset({"127.0.0.1", "::1", "localhost"})
|
||||
|
||||
|
||||
def exempt_from_ceiling(request: Request) -> bool:
|
||||
"""Whether the ceiling should ignore this request.
|
||||
|
||||
Its own function so the rule is testable without standing up a server —
|
||||
and so the healthcheck exemption is somewhere a reader can find it.
|
||||
"""
|
||||
if not request.url.path.startswith("/api/"):
|
||||
return True
|
||||
# The peer address, never the Host header: Host is set by the caller and
|
||||
# would hand every attacker an exemption.
|
||||
return (request.client.host if request.client else "") in _LOCAL_HOSTS
|
||||
|
||||
|
||||
class GlobalRateLimitMiddleware(BaseHTTPMiddleware):
|
||||
"""A cap on total /api/ traffic, independent of any client identity."""
|
||||
|
||||
async def dispatch(self, request: Request, call_next):
|
||||
global _global_window
|
||||
|
||||
if exempt_from_ceiling(request):
|
||||
return await call_next(request)
|
||||
|
||||
now = time.monotonic()
|
||||
# One event loop, and no await between the read and the write, so this
|
||||
# sequence is atomic without a lock.
|
||||
if _global_window is None or now - _global_window[0] >= 60:
|
||||
_global_window = [now, 0]
|
||||
_global_window[1] += 1
|
||||
|
||||
if _global_window[1] > settings.global_rate_limit_per_minute:
|
||||
return JSONResponse(
|
||||
# Distinguishable from slowapi's per-client 429: an operator
|
||||
# reading logs has to be able to tell "one noisy client" from
|
||||
# "the origin is saturated".
|
||||
{"detail": "The service is at capacity. Please retry shortly."},
|
||||
status_code=429,
|
||||
headers={"Retry-After":
|
||||
str(max(1, int(60 - (now - _global_window[0]))))},
|
||||
)
|
||||
return await call_next(request)
|
||||
|
||||
|
||||
class SecurityHeadersMiddleware(BaseHTTPMiddleware):
|
||||
"""Add security headers to all responses."""
|
||||
|
||||
@@ -532,6 +598,10 @@ app.add_middleware(CacheAndETagMiddleware)
|
||||
app.add_middleware(SecurityHeadersMiddleware)
|
||||
app.add_middleware(RequestSizeLimitMiddleware)
|
||||
app.add_middleware(GZipMiddleware, minimum_size=512)
|
||||
# Added last, so it is outermost and refuses before anything downstream does
|
||||
# work. A ceiling that only applies after the expensive part has run is not a
|
||||
# ceiling.
|
||||
app.add_middleware(GlobalRateLimitMiddleware)
|
||||
|
||||
# CORS middleware - restricted for production
|
||||
app.add_middleware(
|
||||
|
||||
@@ -35,6 +35,11 @@ class Settings(BaseSettings):
|
||||
# Security
|
||||
admin_api_key: str = Field(default_factory=lambda: secrets.token_urlsafe(32))
|
||||
rate_limit_per_minute: int = 60 # Requests per minute per IP
|
||||
# A ceiling on total /api/ traffic, independent of any client identity.
|
||||
# client_key trusts headers only Cloudflare can vouch for, so a caller
|
||||
# reaching the origin directly could otherwise mint a fresh bucket per
|
||||
# request. See GlobalRateLimitMiddleware in backend/app.py.
|
||||
global_rate_limit_per_minute: int = 3000
|
||||
rate_limit_burst: int = 10 # Allow burst of requests
|
||||
max_request_size: int = 1024 * 1024 # 1MB max request size
|
||||
|
||||
|
||||
+15
-3
@@ -132,9 +132,21 @@ def suggest_schools_typesense(query: str, limit: int = 8) -> List[dict]:
|
||||
return []
|
||||
|
||||
rows = []
|
||||
for hit in result.get("hits", []):
|
||||
doc = hit.get("document", {})
|
||||
row = {"urn": int(doc.get("urn", 0))}
|
||||
for hit in result.get("hits", []) or []:
|
||||
doc = (hit or {}).get("document") or {}
|
||||
try:
|
||||
urn = int(doc["urn"])
|
||||
except (KeyError, TypeError, ValueError):
|
||||
# Skip the row, keep the rest. Typesense declares urn as int32 so
|
||||
# this should be unreachable, but the index is a separate system
|
||||
# that something other than this code can reindex — and "never
|
||||
# raises" is a promise the keystroke path actually depends on.
|
||||
# Dropping one malformed document is right; blanking the whole
|
||||
# dropdown, or serving a suggestion pointing at /school/0, is not.
|
||||
logging.getLogger(__name__).warning(
|
||||
"skipping malformed suggestion document: %r", doc)
|
||||
continue
|
||||
row = {"urn": urn}
|
||||
row.update({f: str(doc.get(f, "") or "") for f in _SUGGEST_FIELDS})
|
||||
rows.append(row)
|
||||
return rows
|
||||
|
||||
@@ -56,3 +56,92 @@ def test_whitespace_is_stripped():
|
||||
# first; an unstripped key silently creates a second bucket per client.
|
||||
assert client_key(_Req({"x-forwarded-for": " 203.0.113.7 ,10.0.0.2"})) \
|
||||
== "203.0.113.7"
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# The ceiling that header rotation cannot raise.
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
import pytest
|
||||
from fastapi.testclient import TestClient
|
||||
|
||||
|
||||
@pytest.fixture()
|
||||
def api(monkeypatch):
|
||||
from backend import app as app_module
|
||||
from backend.config import settings
|
||||
|
||||
monkeypatch.setattr(settings, "global_rate_limit_per_minute", 5)
|
||||
monkeypatch.setattr(app_module, "_global_window", None)
|
||||
return TestClient(app_module.app, raise_server_exceptions=False)
|
||||
|
||||
|
||||
|
||||
def _ceiling_req(path: str, host: str):
|
||||
"""Enough of a Request for exempt_from_ceiling: a path and a peer host."""
|
||||
return type("R", (), {
|
||||
"url": type("U", (), {"path": path})(),
|
||||
"client": type("C", (), {"host": host})(),
|
||||
})()
|
||||
|
||||
|
||||
def _get(client, path="/api/flags", cf=None):
|
||||
headers = {"cf-connecting-ip": cf} if cf else {}
|
||||
return client.get(path, headers=headers)
|
||||
|
||||
|
||||
def test_rotating_the_cloudflare_header_cannot_buy_unlimited_requests(api):
|
||||
"""The attack the per-client keying opened up.
|
||||
|
||||
client_key trusts CF-Connecting-IP, and nothing in this process can tell an
|
||||
edge-set header from an attacker-set one — that distinction can only be
|
||||
made at Cloudflare, with Authenticated Origin Pulls or an origin firewall.
|
||||
A caller reaching the origin directly can therefore mint a fresh
|
||||
rate-limit bucket per request and evade per-client limits entirely.
|
||||
|
||||
Per-client fairness is still the right default; this is the backstop that
|
||||
bounds what evading it can achieve. Without it, correct keying would be a
|
||||
net regression against abuse compared with the shared bucket it replaced.
|
||||
"""
|
||||
codes = [_get(api, cf=f"203.0.113.{i}").status_code for i in range(8)]
|
||||
assert codes.count(200) == 5
|
||||
assert codes.count(429) == 3
|
||||
|
||||
|
||||
def test_the_ceiling_says_which_limit_was_hit(api):
|
||||
# Distinguishable from slowapi's per-client 429, or an operator reading
|
||||
# logs cannot tell "one noisy client" from "the origin is saturated".
|
||||
for i in range(5):
|
||||
_get(api, cf=f"203.0.113.{i}")
|
||||
refused = _get(api, cf="203.0.113.99")
|
||||
assert refused.status_code == 429
|
||||
assert "capacity" in refused.json()["detail"].lower()
|
||||
assert refused.headers.get("retry-after")
|
||||
|
||||
|
||||
def test_traffic_below_the_ceiling_is_untouched(api):
|
||||
codes = [_get(api, cf=f"203.0.113.{i}").status_code for i in range(5)]
|
||||
assert codes == [200] * 5
|
||||
|
||||
|
||||
def test_the_container_healthcheck_is_exempt(api):
|
||||
"""The healthcheck runs `curl http://localhost:80/api/data-info` inside the
|
||||
container. If the ceiling could starve it, saturation would fail the
|
||||
healthcheck, restart the container, and turn a load spike into an outage
|
||||
loop — the ceiling has to protect the origin, not kill it.
|
||||
"""
|
||||
from backend.app import exempt_from_ceiling
|
||||
|
||||
assert exempt_from_ceiling(_ceiling_req("/api/data-info", "127.0.0.1"))
|
||||
assert exempt_from_ceiling(_ceiling_req("/api/data-info", "::1"))
|
||||
# Everyone else is counted.
|
||||
assert not exempt_from_ceiling(_ceiling_req("/api/data-info", "10.0.0.9"))
|
||||
|
||||
|
||||
def test_the_ceiling_ignores_non_api_paths():
|
||||
# Sitemaps and robots.txt are served by this app too, and a crawler
|
||||
# fetching them must not be refused because the API is busy.
|
||||
from backend.app import exempt_from_ceiling
|
||||
|
||||
assert exempt_from_ceiling(_ceiling_req("/sitemap.xml", "10.0.0.9"))
|
||||
assert exempt_from_ceiling(_ceiling_req("/robots.txt", "10.0.0.9"))
|
||||
@@ -126,3 +126,32 @@ def test_the_response_is_cacheable(monkeypatch):
|
||||
res = _client(monkeypatch, [_HIT]).get("/api/suggest?q=breck")
|
||||
assert "s-maxage" in res.headers.get("cache-control", "")
|
||||
assert res.headers.get("etag")
|
||||
|
||||
|
||||
def test_a_malformed_urn_does_not_raise(monkeypatch):
|
||||
"""The docstring promises "never raises"; the parsing loop sat outside the
|
||||
try, so int(None) or int("abc") would have turned a keystroke into a 500.
|
||||
|
||||
Typesense declares urn as int32, so this should be unreachable — but the
|
||||
contract is what the caller relies on, and a search index is a separate
|
||||
system that can be reindexed by something other than this code.
|
||||
"""
|
||||
_use(monkeypatch, _FakeClient([{"urn": None, "school_name": "X",
|
||||
"local_authority": "Y", "postcode": "Z"}]))
|
||||
assert data_loader.suggest_schools_typesense("x") == []
|
||||
|
||||
|
||||
def test_a_malformed_row_does_not_discard_the_good_ones(monkeypatch):
|
||||
# One bad document must not blank the whole dropdown.
|
||||
_use(monkeypatch, _FakeClient([
|
||||
{"urn": "not-a-number", "school_name": "Bad", "local_authority": "Y",
|
||||
"postcode": "Z"},
|
||||
_HIT,
|
||||
]))
|
||||
out = data_loader.suggest_schools_typesense("x")
|
||||
assert [r["urn"] for r in out] == [100010]
|
||||
|
||||
|
||||
def test_a_hit_with_no_document_does_not_raise(monkeypatch):
|
||||
_use(monkeypatch, _FakeClient([{}]))
|
||||
assert data_loader.suggest_schools_typesense("x") == []
|
||||
Reference in new issue
Block a user