feat(suggest): GET /api/suggest, cacheable and DataFrame-free

A dedicated endpoint rather than a mode of /api/schools, because that
path filters and sorts 25,000 pandas rows per query while holding the
GIL — affordable once per search, not once per keystroke. A test asserts
the distinction directly by making load_school_data raise and requiring
the endpoint to answer anyway.

Nothing errors on ordinary input: a short query, no matches, or
Typesense being down are all 200 with an empty list.

Cached deliberately. Prefix queries repeat enormously across users and
school names change once a year, so s-maxage plus the existing ETag
middleware turns most keystrokes into 304s.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
This commit is contained in:
TudorandClaude Opus 5 committed 2026-08-26 20:34:34 +01:00
1 parent 75d3534d82
commit 1a6d349dad
2 files changed
+91

No files matched your search

+34
View File
@@ -33,6 +33,7 @@ from .data_loader import (
get_supplementary_data,
get_supplementary_data_batch,
search_schools_typesense,
suggest_schools_typesense,
)
from .data_loader import get_data_info as get_db_info
from . import flags
@@ -383,6 +384,7 @@ CACHE_RULES: list[tuple[str, tuple[int, int, int]]] = [
("/api/schools/", (300, 3600, 86400)), # /api/schools/{urn}
("/api/rankings", (60, 600, 3600)),
("/api/compare", (60, 600, 3600)),
("/api/suggest", (60, 3600, 86400)), # autosuggest
("/api/schools", (30, 300, 1800)), # search list
]
@@ -1299,6 +1301,38 @@ async def get_place(request: Request, kind: str, slug: str,
}
# Two characters. One is not a query — it matches thousands of schools and the
# response is useless, so it is not worth a round trip.
SUGGEST_MIN_QUERY = 2
@app.get("/api/suggest")
@limiter.limit("120/minute")
async def suggest_schools(
request: Request,
q: str = Query("", max_length=100),
limit: int = Query(8, ge=1, le=20),
):
"""School name suggestions, from Typesense alone.
Deliberately not a mode of /api/schools: that path filters and sorts the
full in-memory DataFrame, which is far too expensive to run per keystroke.
Nothing here returns an error for ordinary input. A short query, no
matches, or Typesense being unreachable are all 200 with an empty list —
a dropdown that quietly does not appear is the right failure for a
keystroke path, and there is no DataFrame fallback because the 25,000-row
substring scan is precisely what this endpoint exists to avoid.
120/minute rather than the default 60: a 200 ms debounce makes typing
legitimately bursty.
"""
query = q.strip()
if len(query) < SUGGEST_MIN_QUERY:
return {"suggestions": []}
return {"suggestions": suggest_schools_typesense(query, limit)}
@app.get("/api/flags")
@limiter.limit(f"{settings.rate_limit_per_minute}/minute")
async def get_feature_flags(request: Request):