docs: revise the spec to the design that survived staging
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m15s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 19s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m18s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 11s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 57s

The spec described the tier system as the design of record. It is gone, so
the document was describing something the code deliberately does not do.

The revision note and the "why not, having built it the other way first"
passage are kept rather than overwritten. The mistake is the instructive part:
treating a preference as a constraint inverted the ranking, and the stopping
rule added to prevent weak distant matches is what guaranteed six Catholic
schools and no community school down the road. A spec that quietly presents the
second design as the plan teaches nobody why the first one failed.

The mockup link is annotated as one revision behind rather than silently left
to look current.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
TudorandClaude Opus 5 committed 2026-09-22 13:06:32 +01:00
1 parent cd6a45bf7d
commit cd1c5d1e1a
1 file changed
+101 -86
@@ -1,9 +1,16 @@
# Similar Schools Nearby — Design
# Other Schools Nearby — Design
**Date:** 2026-09-21
**Status:** awaiting review
**Date:** 2026-09-21, revised 2026-09-22
**Status:** revised after staging review
**Scope:** school detail pages, both phase templates
> **Revision, 2026-09-22.** The first build ranked by intake similarity and used
> distance as a tiebreak. On staging a Catholic primary showed six Catholic
> primaries, none of them close enough to be a real option, and omitted the
> community school down the road. Distance now decides the order and nothing
> else does; the tier system is gone. The reasoning is kept below rather than
> quietly overwritten, because the mistake is the instructive part.
## Goal
Give a school detail page an answer to the question every reader arrives with
@@ -15,7 +22,8 @@ school. This section adds that edge — up to six nearby schools of the same pha
and a comparable intake, three at a time in a carousel, each a crawlable link and
each addable to the comparison basket in one click.
Mockup, with all three tier states live in both themes:
Mockup, in both themes (drawn against the original tiered design, so its ledes
and chip fallbacks are one revision behind the copy specified below):
<https://claude.ai/artifact/168KdUMcfkUeGWW2FGjuec>
Source of the same page in the repo: `mockups/similar-schools-nearby.html`.
@@ -37,19 +45,23 @@ alternative, and there are three ways that claim goes wrong:
3. A **single-sex** school of the opposite sex. Not a weak match — not an option
at all.
So the design separates two kinds of rule, and never confuses them:
So the design separates two kinds of fact, and never confuses them:
- **Hard filters** encode the claims above. They are never relaxed, at any
distance, even if that means the section does not render.
- **Soft preferences** describe how closely the intake resembles this school's.
They relax in tiers, and the card's own text always states what survived.
- **Hard filters** encode the claims above. They decide eligibility, and are
never relaxed at any distance, even if that means the section does not render.
- **Shared characteristics** — gender, religious character, selectivity —
describe how closely an intake resembles this school's. They are *reported on
the card and never ranked on*, so the reader weighs them rather than having
them weighed for them.
Everything below follows from that split.
Everything below follows from that split. The revision at the top of this
document is what happens when the second kind is treated as the first.
## Selection algorithm
A backend helper, `_similar_schools_payload(urn)` in `backend/app.py`, modelled
on the existing `_places_payload(urn)` and operating on the cached
A backend helper, `_nearby_schools_payload(urn)` in `backend/app.py`, modelled
on the existing `_places_payload(urn)` and delegating to
`backend/nearby_schools.select_nearby(frame, urn)`, which operates on the cached
`load_latest_school_data()` frame — one row per URN, already carrying
`latitude`, `longitude`, `phase`, `gender`, `religious_denomination`,
`admissions_policy`, `school_type` and `status`.
@@ -77,64 +89,63 @@ school the frontend treats as special for benchmarking but the backend treats as
mainstream for matching would be dropped from its own England comparison and
then offered as a peer to a mainstream school on the next page along.
**Up to six cards, three visible.** Six is a cap, not a quota: the section shows
every school that qualifies at the tiers it used, up to six. Three fit the row,
and the rest are reached with the carousel arrows. Two is the minimum that
renders at all.
**Up to six cards, three visible.** Three fit the row; the rest are reached with
the carousel arrows. Two is the minimum that renders at all.
### Soft preferences, relaxed in tiers
### Order: distance, and nothing else
| Tier | Additionally requires | Radius |
|---|---|---|
| 1 | exact gender equality **and** same religious character | 3 miles |
| 2 | exact gender equality | 5 miles |
| 3 | nothing beyond the hard filters | 10 miles |
The nearest eligible schools, closest first. Similarity does not enter the
ranking at any point.
**Tiers relax to reach a usable set, never to fill the last slots.**
**Why not, having built it the other way first.** The original design ranked by
tiers — same gender and faith within 3 miles, then same gender within 5, then
anything within 10 — and used distance only to order the result. Two things
followed, and both showed up on the first Catholic primary anyone looked at:
Work down the tiers until the schools found so far reach three. Call the tier
that got there T. The section then shows up to six schools drawn from tiers 1
to T, nearest first — and does not open tier T+1 merely because six slots are
not yet full.
- A faith match at 2.9 miles outranked a community school at 0.3 miles. For a
primary, whose catchment is routinely under a mile, the far school is not a
weaker option; it is not an option.
- Because the row filled from the best tier before widening, three Catholic
schools within 3 miles were enough to fill all six slots with Catholic
schools. The stopping rule that produced this had been added to prevent the
*opposite* failure — padding a row with weak distant matches — and made this
one certain.
Worked through:
The premise was backwards. **Distance is a constraint and intake is a
preference.** A parent cannot act on a school outside their reach however well
it matches, and they are perfectly capable of noticing a shared denomination
for themselves if we show it to them. So similarity moved from the ranking to
the card: `shared` reports what a school genuinely has in common, and the reader
applies their own weighting.
| Qualifying | T | Shown |
|---|---|---|
| 14 at tier 1 | 1 | the 6 nearest tier-1 schools |
| 4 at tier 1 | 1 | all 4 — tier 2 is never opened |
| 2 at tier 1, 7 more at tier 2 | 2 | the 6 nearest of those 9 |
| 2 at tier 1, 1 at tier 2 | 2 | all 3 |
| 2 across all three tiers | 3 | both, since 2 is the minimum |
The hard filters above were always where the defensibility lived. They are
untouched.
Without that stopping rule, a cap of six would reliably drag in tier-3 schools
ten miles away to fill a row that three good matches had already earned. The old
cap of three hid this; six exposes it, which is why the rule is stated rather
than left to the loop.
### Reach: a sanity bound, not a target
**Faith relaxes before gender.** A faith mismatch changes the character of a
school; a gender mismatch can mean the school is not available to the reader's
child at all. Ordering them the other way would fill the section with schools
that cannot be applied to.
| Phase | Reach |
|---|---|
| Primary, middle deemed primary, all-through | 2 miles |
| Secondary, middle deemed secondary | 6 miles |
| 16 plus | 10 miles |
### Two decisions that are easy to get wrong later
Ordering by distance already handles density — a school in inner London fills
all six slots inside a mile and never approaches the cap. The cap decides one
thing: what happens where the area is sparse. It differs by phase because
catchments do, and because people travel furthest for post-16.
**Selected by tier, displayed by distance.** Tier decides *which* three schools
earn a slot. The rendered order is then distance ascending, because "nearby" is
the promise in the heading and a reader scanning the row reads the first card as
the closest. A tier-2 school at 0.4 miles therefore appears above a tier-1
school at 2.9 miles, and the chips explain the difference in match quality.
**A primary with nothing inside two miles renders no section**, and that is the
intended answer rather than a gap. The alternative is a section headed "nearby"
listing a school four miles from a five-year-old.
**Past the sixth school, the rest are dropped without a count.** In inner
London dozens clear tier 1, and a parent there will notice three is not the
neighbourhood — hence six. Beyond that the section does not try to be the list:
`NearbyPlaces` sits directly beneath and already leads to the place pages, which
are built for browsing a full set and which the school page exists to feed.
**Past the sixth school, the rest are dropped without a count.** The section
does not try to be the list: `NearbyPlaces` sits directly beneath and already
leads to the place pages, which are built for browsing a full set and which the
school page exists to feed.
**Fewer than two results renders nothing.** Not an empty state, not a single
lonely card, not padding with schools that failed the hard filters. The section
is absent, the nav item is absent, and the page is unchanged from today. A page
with one weak match is better off without the section than with it.
lonely card. The section is absent, the nav item is absent, and the page is
unchanged from before it existed.
### Distance
@@ -157,14 +168,13 @@ reader to assume otherwise.
| `distance_miles` | one decimal place |
| `school_type` | GIAS type, translated, for the card's meta line |
| `age_range` | for the meta line |
| `shared` | the chip strings the tier actually justifies — see below |
| `tier` | 1, 2 or 3 — drives the lede's wording and the chip styling |
| `shared` | what this school genuinely shares with the subject; may be empty |
| `metric_value` | the phase-appropriate headline figure, or null |
| `metric_key` | `rwm_expected_pct` or `attainment_8_score` — see below |
| `metric_year` | the year the figure is from |
The metric follows the phase side the school was *matched* on, not the
neighbour's own phase, so a row of cards never mixes two scales. The secondary
The metric follows the subject school's phase side, not the neighbour's own
phase, so a row of cards never mixes two scales. The secondary
side uses `attainment_8_score`; the primary side uses `rwm_expected_pct`. Where
the neighbour has no value for that key, the card reads "Not published" rather
than falling back to the other key.
@@ -183,12 +193,6 @@ correctly — against secondaries. The section therefore takes its lede noun fro
the school's own phase rather than from its template, or it would print "Other
primary schools near <sixth form college>" above a row of secondaries.
`tier` is carried explicitly rather than inferred from the contents of
`shared`, because the frontend needs it for two separate decisions — whether the
lede may claim a similar intake, and whether a chip renders as a brand-tinted
fill or a muted outline — and inferring it from chip count would couple those
decisions to the copy.
Up to six rows of roughly 130 bytes each. It rides in the existing detail payload
rather than a new endpoint because the page already makes exactly one server
fetch for its data, and `/school/[slug]` regenerates at most weekly
@@ -199,9 +203,11 @@ week.
frontend treats absent and empty identically, which is what allowed
`NearbyPlaces` to ship without a lockstep deploy of the two images.
`shared` is computed on the backend beside the tier that produced it, not
re-derived on the frontend. Deriving it twice is how a card comes to claim a
match the selection did not actually make.
`shared` is computed on the backend, beside the data it is derived from, not
re-derived on the frontend. Deriving it twice is how a card comes to claim
something the selection never established. An empty list is a real answer and
renders no chips: a bare card costs a school nothing but the likeness it does
not have, since the order was already settled by distance.
## Frontend
@@ -308,16 +314,21 @@ shows.
## Copy, and what the section is allowed to claim
**The lede tracks the deepest tier shown.** At tiers 1–2 it reads "Other primary
schools near X, with a similar intake." Where any card came from tier 3 it drops
"with a similar intake", because for at least one of the cards that is not what
was matched. Six cards make this more likely to fire than three did, which is
correct: a wider net is exactly when the claim needs dropping.
**The lede never claims an intake.** It reads "Other primary schools near X." —
one sentence, no variants. The earlier version varied the wording by tier, which
only existed to soften a claim the section should not have been making.
**Chips state only what is shared.** A tier-2 card carries fewer chips rather
than a chip it has not earned; a tier-3 card falls back to the plain phase name,
styled as a muted outline rather than a brand-tinted fill so the difference is
visible at a glance.
**The heading is "Other schools nearby", not "Similar schools nearby".** The
hard filters do guarantee a comparable set — same phase, same selectivity,
mainstream never beside special — but nothing ranks on likeness, so the heading
does not say it does. The nav item reads "Nearby schools" and the section id is
`nearby`.
**Chips state only what is shared, and may be absent entirely.** A card with
nothing in common renders no chip row rather than falling back to a filler.
Since chips no longer affect the order, an empty one costs that school nothing
except a claim it cannot support — and a Catholic parent scanning the row still
spots "Roman Catholic" on the card that carries it, and weighs it themselves.
**The neighbour's metric carries no valence colour.** Green and terracotta are
reserved site-wide for comparison against the England average. Colouring a
@@ -364,9 +375,11 @@ synthetic frame rather than live marts:
- a special school returns only special schools; a mainstream school returns none
- a Boys school never returns a Girls school; Mixed matches both
- closed schools and schools without coordinates are never returned
- tier relaxation fills in order, and a school taken at tier 1 is not repeated
- tiers stop relaxing once three are found: four tier-1 matches never open tier 2
- results are ordered by distance ascending, always
- a faith match never outranks a closer school (the staging defect, pinned)
- the nearest eligible school is always present
- more than six qualifying schools returns the six nearest
- reach is capped per phase, and a primary beyond two miles returns `[]`
- an all-through school is offered on both phase sides
- a `16 plus` school is matched against secondaries and colleges, never primaries
- `is_secondary_phase` and `PHASE_GROUPS` agree on every GIAS phase value
@@ -376,7 +389,8 @@ synthetic frame rather than live marts:
**Frontend**, in `nextjs-app/__tests__`:
- the section renders nothing for absent, empty and single-row inputs
- the lede drops "with a similar intake" when any card is tier 3
- the lede never claims a similar intake
- an empty `shared` renders no chips rather than a filler
- a null metric renders "Not published"
- the nav item appears only alongside the section
- every card is in the DOM, including the ones scrolled out of view
@@ -408,10 +422,11 @@ journeys are confirmed on the post-merge staging run.
- Autoplay, dots, or an infinite loop on the carousel. It is a short list a
reader scans deliberately, not a banner competing for attention, and a row
that moves on its own is a row that moves while someone is reading it.
- Statistical neighbours on deprivation, size or cohort profile. If the tiers
prove too coarse, that is the trigger to move this computation into a dbt mart
— `_similar_schools_payload` is a deliberate seam for exactly that swap.
- Statistical neighbours on deprivation, size or cohort profile. If plain
distance proves too blunt, that is the trigger to move this computation into a
dbt mart — `select_nearby` is a deliberate seam for exactly that swap.
- Precomputing neighbours in `marts.*`. Rejected for now: a new mart is inert
until Airflow runs, so the feature would ship dark, and every tuning change to
the tiers would become a pipeline round-trip instead of a deploy.
the rules would become a pipeline round-trip instead of a deploy. The revision
at the top of this document is the argument for keeping that loop short.
- Any change to `/api/compare`, the compare page, or the comparison basket.