fix: order nearby schools by distance, not by how alike they are

Reported from staging: a Catholic primary showed six Catholic primaries, none
of them close enough to be a real option, and omitted the community school
down the road.

Three causes, compounding. Ranking put tier before distance, so a faith match
at 2.9 miles outranked a community school at 0.3. The ENOUGH=3 stopping rule —
added so a cap of six would not drag in weak distant matches — filled the row
from the best tier before it ever widened, which is what made every card
Catholic. And a 3-mile tier-1 radius is sane for a secondary and most of a city
for a primary, whose catchments are routinely under a mile.

The premise was backwards. For a parent, distance is a constraint and intake is
a preference; a school beyond a primary catchment is not a weaker option, it is
not an option. So distance now decides the order and nothing else does. The
hard filters are untouched — they were always where the defensibility lived.
Similarity survives as chips on the card: reported, so a reader applies their
own weighting, rather than ranked, so we apply ours for them.

Reach is capped per phase (primary 2, secondary 6, post-16 10) as a sanity
bound, not a target: ordering already handles density, so the cap only decides
what happens where an area is sparse. A primary with nothing inside two miles
now renders no section, which is the honest answer.

Deleted: the tier system, the stopping rule, the tier-dependent lede, the
`tier` field, the tier-3 fallback chip and its style. select_similar also stops
taking is_secondary — it reads the phase from the subject's own row, so no
caller can hand it one that disagrees with the data.

The heading is now "Other schools nearby". The hard filters still guarantee a
comparable set, but nothing ranks on likeness, so the heading no longer says it
does.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
TudorandClaude Opus 5 committed 2026-09-22 13:06:32 +01:00
1 parent 151cf4bc80
commit 5c0ccc693d
11 files changed
+343 -278

No files matched your search

+79 -81
View File
@@ -1,17 +1,24 @@
"""Which nearby schools a detail page may offer as alternatives.
Two kinds of rule, and they are not interchangeable.
HARD FILTERS decide eligibility, and encode claims the section is not allowed
to make. A selective school is not an alternative to a non-selective one, a
special school is not comparable to a mainstream one, and a Girls school is not
an option for a Boys school's reader. They never relax, at any distance, even
where that means the section does not render at all.
HARD FILTERS encode claims the section is not allowed to make. A selective
school is not an alternative to a non-selective one, a special school is not
comparable to a mainstream one, and a Girls school is not an option for a Boys
school's reader. These never relax, at any distance, even where that means the
section does not render at all.
DISTANCE decides the order, and nothing else does.
SOFT PREFERENCES describe how closely an intake resembles this school's. They
relax in tiers, and every card reports the tier that actually took it so the
page can say what is shared rather than implying more. They relax only far
enough to reach a usable set, never far enough to fill the last of the slots.
An earlier version ranked by intake similarity first and used distance only as
a tiebreak. That put a Catholic school 2.9 miles away above the community
school 0.3 miles down the road, and — because the row filled from the best tier
before widening — filled all six slots with faith matches while omitting every
school a parent could actually walk to. For a primary, a school that far is not
a weaker option; it is not an option. Distance is a constraint and intake is a
preference, and the ranking now says so.
Similarity survives as `shared`: what a candidate genuinely has in common with
this school, reported on its card, so a reader applies their own weighting
instead of having ours applied for them.
Pure functions over a DataFrame: no I/O, no FastAPI, no database.
"""
@@ -25,19 +32,21 @@ import pandas as pd
from .schemas import PHASE_GROUPS
# A cap, not a quota: the section shows everything that qualified at the tiers
# it used, up to this many. Three fit the row; the rest are behind the arrows.
# Three fit the row; the rest are behind the carousel arrows.
MAX_SCHOOLS = 6
# Tiers stop relaxing once this many have been found. Without it, a cap of six
# would reliably drag in tier-3 schools ten miles away to fill a row that three
# good matches had already earned.
ENOUGH = 3
MINIMUM = 2
# (tier, radius in miles). Faith relaxes before gender: a faith mismatch
# changes the character of a school, while a gender mismatch can mean the
# school is not available to this reader's child at all.
TIERS: tuple[tuple[int, float], ...] = ((1, 3.0), (2, 5.0), (3, 10.0))
# How far the section will reach, in miles, when nothing closer exists.
#
# A sanity bound rather than a target: ordering by distance already handles
# density, so a school in a dense area fills all six slots inside a mile and
# never sees this. It decides one thing — what happens where the area is
# sparse — and the answer differs by phase because catchments do. Primary
# catchments are routinely under a mile; beyond two, a primary is not a weaker
# option but not an option, and no section is the honest answer.
PRIMARY_RADIUS_MILES = 2.0
SECONDARY_RADIUS_MILES = 6.0
POST16_RADIUS_MILES = 10.0
EARTH_RADIUS_MILES = 3958.8
@@ -80,15 +89,6 @@ def genders_compatible(a: str | None, b: str | None) -> bool:
return not (left in single and right in single and left != right)
def phase_label(phase: str | None) -> str:
text = (phase or "").strip()
if not text:
return "School"
if text.lower() == "all-through":
return "All-through school"
return f"{text.capitalize()} school"
def is_secondary_phase(phase: str | None) -> bool:
"""Whether this phase takes the secondary side: secondary group membership,
minus all-through.
@@ -107,6 +107,13 @@ def is_secondary_phase(phase: str | None) -> bool:
return text != "all-through" and text in PHASE_GROUPS["secondary"]
def radius_miles(phase: str | None) -> float:
"""How far this phase's section will reach when nothing closer exists."""
if (phase or "").strip().lower() == "16 plus":
return POST16_RADIUS_MILES
return SECONDARY_RADIUS_MILES if is_secondary_phase(phase) else PRIMARY_RADIUS_MILES
def _phase_group(is_secondary: bool) -> set[str]:
return PHASE_GROUPS["secondary" if is_secondary else "primary"]
@@ -144,26 +151,45 @@ def _mask(series: pd.Series, predicate) -> pd.Series:
return pd.Series([predicate(value) for value in series], index=series.index, dtype=bool)
def _chips(subject: pd.Series, candidate: pd.Series, tier: int, is_secondary: bool) -> list[str]:
if tier >= 3:
return [phase_label(candidate.get("phase"))]
def _shared(subject: pd.Series, candidate: pd.Series, is_secondary: bool) -> list[str]:
"""What this candidate genuinely has in common with the subject.
Empty is a real answer, and renders no chips at all. A card claiming a
shared characteristic it does not have would be worse than a bare one —
and since these no longer affect the order, an empty list costs the school
nothing but its place in the row, which distance already decided.
"""
shared: list[str] = []
gender = str(subject.get("gender") or "").strip()
if gender and str(candidate.get("gender") or "").strip().lower() == gender.lower():
shared.append(gender)
chips = [str(subject.get("gender") or "").strip()]
if is_secondary:
policy = (candidate.get("admissions_policy") or "").strip()
if policy and policy.lower() not in {"not applicable", "unknown"}:
chips.append(policy)
if tier == 1:
chips.append(faith_label(candidate.get("religious_denomination")))
return [chip for chip in chips if chip]
policy = str(candidate.get("admissions_policy") or "").strip()
subject_policy = str(subject.get("admissions_policy") or "").strip()
if (
policy
and policy.lower() == subject_policy.lower()
and policy.lower() not in {"not applicable", "unknown"}
):
shared.append(policy)
if faith_key(candidate.get("religious_denomination")) == faith_key(
subject.get("religious_denomination")
):
shared.append(faith_label(candidate.get("religious_denomination")))
return shared
def select_similar(frame: pd.DataFrame, urn: int, is_secondary: bool) -> list[dict]:
"""Up to MAX_SCHOOLS nearby schools this page may offer, or [] below MINIMUM.
def select_similar(frame: pd.DataFrame, urn: int) -> list[dict]:
"""The nearest eligible schools, closest first — at most MAX_SCHOOLS, and
none at all below MINIMUM.
Selected by tier, displayed by distance: the tier decides which schools
earn a slot, and the render order is then closest-first, because "nearby"
is the promise in the heading.
The phase is read from the subject's own row rather than passed in, so a
caller cannot hand this a phase that disagrees with the data it selects
from.
"""
subject_rows = frame[frame["urn"] == urn]
if subject_rows.empty:
@@ -174,6 +200,9 @@ def select_similar(frame: pd.DataFrame, urn: int, is_secondary: bool) -> list[di
if lat is None or lon is None:
return []
phase = subject.get("phase")
is_secondary = is_secondary_phase(phase)
reach = radius_miles(phase)
metric_key = "attainment_8_score" if is_secondary else "rwm_expected_pct"
candidates = frame[frame["urn"] != urn].copy()
@@ -208,42 +237,12 @@ def select_similar(frame: pd.DataFrame, urn: int, is_secondary: bool) -> list[di
lat, lon, candidates["latitude"].values, candidates["longitude"].values
).round(1)
# ── Soft preferences, in tiers ──────────────────────────────────────
subject_faith = faith_key(subject.get("religious_denomination"))
subject_gender_key = (subject_gender or "").strip().lower()
same_gender = candidates["gender"].fillna("").str.strip().str.lower() == subject_gender_key
same_faith = _mask(
candidates["religious_denomination"], lambda d: faith_key(d) == subject_faith
)
tier_masks = {
1: same_gender & same_faith,
2: same_gender,
3: pd.Series(True, index=candidates.index),
}
# Descend the tiers only until the set reaches ENOUGH. The tier that gets
# there is the last one opened, and the remaining slots up to MAX_SCHOOLS
# are filled from the tiers already used — never by widening again.
picked: dict[int, tuple[int, pd.Series]] = {}
for tier, radius in TIERS:
within = candidates[tier_masks[tier] & (candidates["distance_miles"] <= radius)]
for _, row in within.sort_values("distance_miles").iterrows():
candidate_urn = int(row["urn"])
if candidate_urn in picked:
continue
picked[candidate_urn] = (tier, row)
if len(picked) >= MAX_SCHOOLS:
break
if len(picked) >= ENOUGH:
break
if len(picked) < MINIMUM:
# ── Nearest first, and nothing else has a say ───────────────────────
within = candidates[candidates["distance_miles"] <= reach]
if len(within) < MINIMUM:
return []
selected = sorted(
picked.values(), key=lambda pair: float(pair[1]["distance_miles"])
)[:MAX_SCHOOLS]
selected = within.sort_values(["distance_miles", "urn"]).head(MAX_SCHOOLS)
return [
{
"urn": int(row["urn"]),
@@ -251,11 +250,10 @@ def select_similar(frame: pd.DataFrame, urn: int, is_secondary: bool) -> list[di
"distance_miles": float(row["distance_miles"]),
"school_type": _native(row.get("school_type")),
"age_range": _native(row.get("age_range")),
"shared": _chips(subject, row, tier, is_secondary),
"tier": tier,
"shared": _shared(subject, row, is_secondary),
"metric_value": _native(row.get(metric_key)),
"metric_key": metric_key,
"metric_year": _native(row.get("year")),
}
for tier, row in selected
for _, row in selected.iterrows()
]