fix(api): bound the search candidate set instead of draining Typesense
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m12s
PR Checks / Backend Smoke (pull_request) Successful in 10s
PR Checks / Build Backend (no push) (pull_request) Successful in 17s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m18s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 39s
PR Checks / AI Code Review (Claude) (pull_request) Failing after 6m44s

Fetching every match kept scoped searches correct but left the number of
round trips in the caller's hands: a one-letter query, or a deliberately
broad one, walked the whole collection a page at a time.

Cap the candidate set at 1,000 URNs — four pages — and return the
relevance-ordered prefix when the ceiling is hit. That is still far more
than one page, so the API's own authority, phase and postcode filters
keep the matches they need, while latency and upstream load stay bounded.
A capped query is logged so a genuinely truncated search is visible.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
TudorandClaude Opus 5 committed 2026-09-15 11:21:45 +01:00
1 parent 0c901cd0d1
commit 5ad1cbfb53
3 files changed
+75 -11

No files matched your search

+30 -9
View File
@@ -84,34 +84,55 @@ def _get_typesense_client():
return None
def search_schools_typesense(query: str) -> Optional[List[int]]:
"""Return all matching URNs in relevance order; None means unavailable.
SEARCH_PAGE_SIZE = 250
# Search results are filtered again by the API (authority, phase, postcode,
# etc.), so one page is too small for scoped searches. Keep the candidate set
# bounded, though: a broad query must not turn into an unbounded sequence of
# Typesense requests. Four pages is enough to preserve useful scoped matches
# while putting a hard ceiling on latency and upstream load.
SEARCH_MAX_CANDIDATES = 1_000
Filtering and user pagination happen in the API after this search. Returning
only the first search page would silently discard valid local matches.
Never return a partial candidate set if a later page fails.
def search_schools_typesense(query: str) -> Optional[List[int]]:
"""Return a bounded set of matching URNs in relevance order.
``None`` means Typesense is unavailable; ``[]`` is a valid zero-match
result. The API applies its remaining filters after this search, so the
first few pages are fetched rather than only the first page. Once the
candidate ceiling is reached, the relevance-ordered prefix is returned on
purpose; fetching every match would make common or adversarial queries
unbounded.
"""
client = _get_typesense_client()
if client is None:
return None
urns = []
urns: list[int] = []
fetched = 0
try:
page = 1
while True:
while fetched < SEARCH_MAX_CANDIDATES:
page_size = min(SEARCH_PAGE_SIZE, SEARCH_MAX_CANDIDATES - fetched)
result = client.collections["schools"].documents.search({
"q": query,
"query_by": "school_name,local_authority,postcode",
"per_page": 250,
"per_page": page_size,
"page": page,
"typo_tokens_threshold": 1,
})
hits = result.get("hits", [])
urns.extend(int(h["document"]["urn"]) for h in hits)
if len(urns) >= result.get("found", len(urns)):
fetched += len(hits)
if fetched >= result.get("found", fetched):
return list(dict.fromkeys(urns))
if not hits:
raise ValueError("Search pagination ended before all matches arrived")
page += 1
logging.getLogger(__name__).info(
"Typesense search capped at %d candidates for query %r",
SEARCH_MAX_CANDIDATES,
query,
)
return list(dict.fromkeys(urns))
except Exception:
logging.getLogger(__name__).exception("School search unavailable")
return None