fix: order nearby schools by distance, not by how alike they are
Reported from staging: a Catholic primary showed six Catholic primaries, none of them close enough to be a real option, and omitted the community school down the road. Three causes, compounding. Ranking put tier before distance, so a faith match at 2.9 miles outranked a community school at 0.3. The ENOUGH=3 stopping rule — added so a cap of six would not drag in weak distant matches — filled the row from the best tier before it ever widened, which is what made every card Catholic. And a 3-mile tier-1 radius is sane for a secondary and most of a city for a primary, whose catchments are routinely under a mile. The premise was backwards. For a parent, distance is a constraint and intake is a preference; a school beyond a primary catchment is not a weaker option, it is not an option. So distance now decides the order and nothing else does. The hard filters are untouched — they were always where the defensibility lived. Similarity survives as chips on the card: reported, so a reader applies their own weighting, rather than ranked, so we apply ours for them. Reach is capped per phase (primary 2, secondary 6, post-16 10) as a sanity bound, not a target: ordering already handles density, so the cap only decides what happens where an area is sparse. A primary with nothing inside two miles now renders no section, which is the honest answer. Deleted: the tier system, the stopping rule, the tier-dependent lede, the `tier` field, the tier-3 fallback chip and its style. select_similar also stops taking is_secondary — it reads the phase from the subject's own row, so no caller can hand it one that disagrees with the data. The heading is now "Other schools nearby". The hard filters still guarantee a comparable set, but nothing ranks on likeness, so the heading no longer says it does. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
1 parent
151cf4bc80
commit
5c0ccc693d
11 files changed
+343
-278
No files matched your search
+79
-81
@@ -1,17 +1,24 @@
|
||||
"""Which nearby schools a detail page may offer as alternatives.
|
||||
|
||||
Two kinds of rule, and they are not interchangeable.
|
||||
HARD FILTERS decide eligibility, and encode claims the section is not allowed
|
||||
to make. A selective school is not an alternative to a non-selective one, a
|
||||
special school is not comparable to a mainstream one, and a Girls school is not
|
||||
an option for a Boys school's reader. They never relax, at any distance, even
|
||||
where that means the section does not render at all.
|
||||
|
||||
HARD FILTERS encode claims the section is not allowed to make. A selective
|
||||
school is not an alternative to a non-selective one, a special school is not
|
||||
comparable to a mainstream one, and a Girls school is not an option for a Boys
|
||||
school's reader. These never relax, at any distance, even where that means the
|
||||
section does not render at all.
|
||||
DISTANCE decides the order, and nothing else does.
|
||||
|
||||
SOFT PREFERENCES describe how closely an intake resembles this school's. They
|
||||
relax in tiers, and every card reports the tier that actually took it so the
|
||||
page can say what is shared rather than implying more. They relax only far
|
||||
enough to reach a usable set, never far enough to fill the last of the slots.
|
||||
An earlier version ranked by intake similarity first and used distance only as
|
||||
a tiebreak. That put a Catholic school 2.9 miles away above the community
|
||||
school 0.3 miles down the road, and — because the row filled from the best tier
|
||||
before widening — filled all six slots with faith matches while omitting every
|
||||
school a parent could actually walk to. For a primary, a school that far is not
|
||||
a weaker option; it is not an option. Distance is a constraint and intake is a
|
||||
preference, and the ranking now says so.
|
||||
|
||||
Similarity survives as `shared`: what a candidate genuinely has in common with
|
||||
this school, reported on its card, so a reader applies their own weighting
|
||||
instead of having ours applied for them.
|
||||
|
||||
Pure functions over a DataFrame: no I/O, no FastAPI, no database.
|
||||
"""
|
||||
@@ -25,19 +32,21 @@ import pandas as pd
|
||||
|
||||
from .schemas import PHASE_GROUPS
|
||||
|
||||
# A cap, not a quota: the section shows everything that qualified at the tiers
|
||||
# it used, up to this many. Three fit the row; the rest are behind the arrows.
|
||||
# Three fit the row; the rest are behind the carousel arrows.
|
||||
MAX_SCHOOLS = 6
|
||||
# Tiers stop relaxing once this many have been found. Without it, a cap of six
|
||||
# would reliably drag in tier-3 schools ten miles away to fill a row that three
|
||||
# good matches had already earned.
|
||||
ENOUGH = 3
|
||||
MINIMUM = 2
|
||||
|
||||
# (tier, radius in miles). Faith relaxes before gender: a faith mismatch
|
||||
# changes the character of a school, while a gender mismatch can mean the
|
||||
# school is not available to this reader's child at all.
|
||||
TIERS: tuple[tuple[int, float], ...] = ((1, 3.0), (2, 5.0), (3, 10.0))
|
||||
# How far the section will reach, in miles, when nothing closer exists.
|
||||
#
|
||||
# A sanity bound rather than a target: ordering by distance already handles
|
||||
# density, so a school in a dense area fills all six slots inside a mile and
|
||||
# never sees this. It decides one thing — what happens where the area is
|
||||
# sparse — and the answer differs by phase because catchments do. Primary
|
||||
# catchments are routinely under a mile; beyond two, a primary is not a weaker
|
||||
# option but not an option, and no section is the honest answer.
|
||||
PRIMARY_RADIUS_MILES = 2.0
|
||||
SECONDARY_RADIUS_MILES = 6.0
|
||||
POST16_RADIUS_MILES = 10.0
|
||||
|
||||
EARTH_RADIUS_MILES = 3958.8
|
||||
|
||||
@@ -80,15 +89,6 @@ def genders_compatible(a: str | None, b: str | None) -> bool:
|
||||
return not (left in single and right in single and left != right)
|
||||
|
||||
|
||||
def phase_label(phase: str | None) -> str:
|
||||
text = (phase or "").strip()
|
||||
if not text:
|
||||
return "School"
|
||||
if text.lower() == "all-through":
|
||||
return "All-through school"
|
||||
return f"{text.capitalize()} school"
|
||||
|
||||
|
||||
def is_secondary_phase(phase: str | None) -> bool:
|
||||
"""Whether this phase takes the secondary side: secondary group membership,
|
||||
minus all-through.
|
||||
@@ -107,6 +107,13 @@ def is_secondary_phase(phase: str | None) -> bool:
|
||||
return text != "all-through" and text in PHASE_GROUPS["secondary"]
|
||||
|
||||
|
||||
def radius_miles(phase: str | None) -> float:
|
||||
"""How far this phase's section will reach when nothing closer exists."""
|
||||
if (phase or "").strip().lower() == "16 plus":
|
||||
return POST16_RADIUS_MILES
|
||||
return SECONDARY_RADIUS_MILES if is_secondary_phase(phase) else PRIMARY_RADIUS_MILES
|
||||
|
||||
|
||||
def _phase_group(is_secondary: bool) -> set[str]:
|
||||
return PHASE_GROUPS["secondary" if is_secondary else "primary"]
|
||||
|
||||
@@ -144,26 +151,45 @@ def _mask(series: pd.Series, predicate) -> pd.Series:
|
||||
return pd.Series([predicate(value) for value in series], index=series.index, dtype=bool)
|
||||
|
||||
|
||||
def _chips(subject: pd.Series, candidate: pd.Series, tier: int, is_secondary: bool) -> list[str]:
|
||||
if tier >= 3:
|
||||
return [phase_label(candidate.get("phase"))]
|
||||
def _shared(subject: pd.Series, candidate: pd.Series, is_secondary: bool) -> list[str]:
|
||||
"""What this candidate genuinely has in common with the subject.
|
||||
|
||||
Empty is a real answer, and renders no chips at all. A card claiming a
|
||||
shared characteristic it does not have would be worse than a bare one —
|
||||
and since these no longer affect the order, an empty list costs the school
|
||||
nothing but its place in the row, which distance already decided.
|
||||
"""
|
||||
shared: list[str] = []
|
||||
|
||||
gender = str(subject.get("gender") or "").strip()
|
||||
if gender and str(candidate.get("gender") or "").strip().lower() == gender.lower():
|
||||
shared.append(gender)
|
||||
|
||||
chips = [str(subject.get("gender") or "").strip()]
|
||||
if is_secondary:
|
||||
policy = (candidate.get("admissions_policy") or "").strip()
|
||||
if policy and policy.lower() not in {"not applicable", "unknown"}:
|
||||
chips.append(policy)
|
||||
if tier == 1:
|
||||
chips.append(faith_label(candidate.get("religious_denomination")))
|
||||
return [chip for chip in chips if chip]
|
||||
policy = str(candidate.get("admissions_policy") or "").strip()
|
||||
subject_policy = str(subject.get("admissions_policy") or "").strip()
|
||||
if (
|
||||
policy
|
||||
and policy.lower() == subject_policy.lower()
|
||||
and policy.lower() not in {"not applicable", "unknown"}
|
||||
):
|
||||
shared.append(policy)
|
||||
|
||||
if faith_key(candidate.get("religious_denomination")) == faith_key(
|
||||
subject.get("religious_denomination")
|
||||
):
|
||||
shared.append(faith_label(candidate.get("religious_denomination")))
|
||||
|
||||
return shared
|
||||
|
||||
|
||||
def select_similar(frame: pd.DataFrame, urn: int, is_secondary: bool) -> list[dict]:
|
||||
"""Up to MAX_SCHOOLS nearby schools this page may offer, or [] below MINIMUM.
|
||||
def select_similar(frame: pd.DataFrame, urn: int) -> list[dict]:
|
||||
"""The nearest eligible schools, closest first — at most MAX_SCHOOLS, and
|
||||
none at all below MINIMUM.
|
||||
|
||||
Selected by tier, displayed by distance: the tier decides which schools
|
||||
earn a slot, and the render order is then closest-first, because "nearby"
|
||||
is the promise in the heading.
|
||||
The phase is read from the subject's own row rather than passed in, so a
|
||||
caller cannot hand this a phase that disagrees with the data it selects
|
||||
from.
|
||||
"""
|
||||
subject_rows = frame[frame["urn"] == urn]
|
||||
if subject_rows.empty:
|
||||
@@ -174,6 +200,9 @@ def select_similar(frame: pd.DataFrame, urn: int, is_secondary: bool) -> list[di
|
||||
if lat is None or lon is None:
|
||||
return []
|
||||
|
||||
phase = subject.get("phase")
|
||||
is_secondary = is_secondary_phase(phase)
|
||||
reach = radius_miles(phase)
|
||||
metric_key = "attainment_8_score" if is_secondary else "rwm_expected_pct"
|
||||
|
||||
candidates = frame[frame["urn"] != urn].copy()
|
||||
@@ -208,42 +237,12 @@ def select_similar(frame: pd.DataFrame, urn: int, is_secondary: bool) -> list[di
|
||||
lat, lon, candidates["latitude"].values, candidates["longitude"].values
|
||||
).round(1)
|
||||
|
||||
# ── Soft preferences, in tiers ──────────────────────────────────────
|
||||
subject_faith = faith_key(subject.get("religious_denomination"))
|
||||
subject_gender_key = (subject_gender or "").strip().lower()
|
||||
same_gender = candidates["gender"].fillna("").str.strip().str.lower() == subject_gender_key
|
||||
same_faith = _mask(
|
||||
candidates["religious_denomination"], lambda d: faith_key(d) == subject_faith
|
||||
)
|
||||
|
||||
tier_masks = {
|
||||
1: same_gender & same_faith,
|
||||
2: same_gender,
|
||||
3: pd.Series(True, index=candidates.index),
|
||||
}
|
||||
|
||||
# Descend the tiers only until the set reaches ENOUGH. The tier that gets
|
||||
# there is the last one opened, and the remaining slots up to MAX_SCHOOLS
|
||||
# are filled from the tiers already used — never by widening again.
|
||||
picked: dict[int, tuple[int, pd.Series]] = {}
|
||||
for tier, radius in TIERS:
|
||||
within = candidates[tier_masks[tier] & (candidates["distance_miles"] <= radius)]
|
||||
for _, row in within.sort_values("distance_miles").iterrows():
|
||||
candidate_urn = int(row["urn"])
|
||||
if candidate_urn in picked:
|
||||
continue
|
||||
picked[candidate_urn] = (tier, row)
|
||||
if len(picked) >= MAX_SCHOOLS:
|
||||
break
|
||||
if len(picked) >= ENOUGH:
|
||||
break
|
||||
|
||||
if len(picked) < MINIMUM:
|
||||
# ── Nearest first, and nothing else has a say ───────────────────────
|
||||
within = candidates[candidates["distance_miles"] <= reach]
|
||||
if len(within) < MINIMUM:
|
||||
return []
|
||||
|
||||
selected = sorted(
|
||||
picked.values(), key=lambda pair: float(pair[1]["distance_miles"])
|
||||
)[:MAX_SCHOOLS]
|
||||
selected = within.sort_values(["distance_miles", "urn"]).head(MAX_SCHOOLS)
|
||||
return [
|
||||
{
|
||||
"urn": int(row["urn"]),
|
||||
@@ -251,11 +250,10 @@ def select_similar(frame: pd.DataFrame, urn: int, is_secondary: bool) -> list[di
|
||||
"distance_miles": float(row["distance_miles"]),
|
||||
"school_type": _native(row.get("school_type")),
|
||||
"age_range": _native(row.get("age_range")),
|
||||
"shared": _chips(subject, row, tier, is_secondary),
|
||||
"tier": tier,
|
||||
"shared": _shared(subject, row, is_secondary),
|
||||
"metric_value": _native(row.get(metric_key)),
|
||||
"metric_key": metric_key,
|
||||
"metric_year": _native(row.get("year")),
|
||||
}
|
||||
for tier, row in selected
|
||||
for _, row in selected.iterrows()
|
||||
]
|
||||
Reference in new issue
Block a user