Files
school_compare/docs/superpowers/specs/2026-09-21-similar-schools-nearby-design.md
T

402 lines
19 KiB
Markdown
Raw Normal View History

# Similar Schools Nearby — Design
**Date:** 2026-09-21
**Status:** awaiting review
**Scope:** school detail pages, both phase templates
## Goal
Give a school detail page an answer to the question every reader arrives with
after the results tables: *and what else is around here?*
Today a school page links outward to its place pages through
`components/school/NearbyPlaces.tsx` and nowhere else. It never links to another
school. This section adds that edge — up to six nearby schools of the same phase
and a comparable intake, three at a time in a carousel, each a crawlable link and
each addable to the comparison basket in one click.
Mockup, with all three tier states live in both themes:
<https://claude.ai/artifact/168KdUMcfkUeGWW2FGjuec>
Source of the same page in the repo: `mockups/similar-schools-nearby.html`.
## The constraint that shapes everything
**A nearby school is not automatically a comparable school.**
The section's whole value is that a reader treats what it shows as a shortlist.
That makes every card an implicit claim that the school is a realistic
alternative, and there are three ways that claim goes wrong:
1. A **selective** school beside a non-selective one. Their intakes are
different by construction, so putting their Attainment 8 figures side by side
invites a conclusion the data cannot support.
2. A **special school, PRU or AP** beside a mainstream school. This is the same
error PR #70 fixed for the England benchmark, where Greenmead (URN 101099)
rendered "0% — 62 below England".
3. A **single-sex** school of the opposite sex. Not a weak match — not an option
at all.
So the design separates two kinds of rule, and never confuses them:
- **Hard filters** encode the claims above. They are never relaxed, at any
distance, even if that means the section does not render.
- **Soft preferences** describe how closely the intake resembles this school's.
They relax in tiers, and the card's own text always states what survived.
Everything below follows from that split.
## Selection algorithm
A backend helper, `_similar_schools_payload(urn)` in `backend/app.py`, modelled
on the existing `_places_payload(urn)` and operating on the cached
`load_latest_school_data()` frame — one row per URN, already carrying
`latitude`, `longitude`, `phase`, `gender`, `religious_denomination`,
`admissions_policy`, `school_type` and `status`.
### Hard filters
| Filter | Rule |
|---|---|
| Self | `urn` is excluded |
| Status | GIAS status must be open |
| Coordinates | both `latitude` and `longitude` present on both schools |
| Phase | same phase group via the existing `PHASE_GROUPS` map |
| Provision | special/PRU/AP match only each other |
| Selectivity | selective matches selective; non-selective matches non-selective |
| Gender | Boys never matches Girls; Mixed is compatible with both |
`PHASE_GROUPS` is reused rather than re-derived so an all-through school is
offered correctly on both the primary and secondary sides, exactly as it already
behaves in search.
The provision filter needs a backend counterpart to the frontend's
`isSpecialSchool()` in `nextjs-app/lib/utils.ts:897`, reading the same GIAS
establishment types through `backend/gias_codes.py`. The two must agree: a
school the frontend treats as special for benchmarking but the backend treats as
mainstream for matching would be dropped from its own England comparison and
then offered as a peer to a mainstream school on the next page along.
**Up to six cards, three visible.** Six is a cap, not a quota: the section shows
every school that qualifies at the tiers it used, up to six. Three fit the row,
and the rest are reached with the carousel arrows. Two is the minimum that
renders at all.
### Soft preferences, relaxed in tiers
| Tier | Additionally requires | Radius |
|---|---|---|
| 1 | exact gender equality **and** same religious character | 3 miles |
| 2 | exact gender equality | 5 miles |
| 3 | nothing beyond the hard filters | 10 miles |
**Tiers relax to reach a usable set, never to fill the last slots.**
Work down the tiers until the schools found so far reach three. Call the tier
that got there T. The section then shows up to six schools drawn from tiers 1
to T, nearest first — and does not open tier T+1 merely because six slots are
not yet full.
Worked through:
| Qualifying | T | Shown |
|---|---|---|
| 14 at tier 1 | 1 | the 6 nearest tier-1 schools |
| 4 at tier 1 | 1 | all 4 — tier 2 is never opened |
| 2 at tier 1, 7 more at tier 2 | 2 | the 6 nearest of those 9 |
| 2 at tier 1, 1 at tier 2 | 2 | all 3 |
| 2 across all three tiers | 3 | both, since 2 is the minimum |
Without that stopping rule, a cap of six would reliably drag in tier-3 schools
ten miles away to fill a row that three good matches had already earned. The old
cap of three hid this; six exposes it, which is why the rule is stated rather
than left to the loop.
**Faith relaxes before gender.** A faith mismatch changes the character of a
school; a gender mismatch can mean the school is not available to the reader's
child at all. Ordering them the other way would fill the section with schools
that cannot be applied to.
### Two decisions that are easy to get wrong later
**Selected by tier, displayed by distance.** Tier decides *which* three schools
earn a slot. The rendered order is then distance ascending, because "nearby" is
the promise in the heading and a reader scanning the row reads the first card as
the closest. A tier-2 school at 0.4 miles therefore appears above a tier-1
school at 2.9 miles, and the chips explain the difference in match quality.
**Past the sixth school, the rest are dropped without a count.** In inner
London dozens clear tier 1, and a parent there will notice three is not the
neighbourhood — hence six. Beyond that the section does not try to be the list:
`NearbyPlaces` sits directly beneath and already leads to the place pages, which
are built for browsing a full set and which the school page exists to feed.
**Fewer than two results renders nothing.** Not an empty state, not a single
lonely card, not padding with schools that failed the hard filters. The section
is absent, the nav item is absent, and the page is unchanged from today. A page
with one weak match is better off without the section than with it.
### Distance
Straight-line, from the vectorised haversine already used for postcode search at
`backend/app.py:831`, computed over the ~27k-row frame in numpy. Reported to one
decimal place in miles, consistent with the rest of the site.
Straight-line distance is not road distance and is not measured from the
reader's home. The section says so in its disclosure rather than leaving the
reader to assume otherwise.
## API
`/api/schools/{urn}` gains a `similar_schools` array. Each row:
| Field | Notes |
|---|---|
| `urn` | for the link and the compare basket |
| `school_name` | link text |
| `distance_miles` | one decimal place |
| `school_type` | GIAS type, translated, for the card's meta line |
| `age_range` | for the meta line |
| `shared` | the chip strings the tier actually justifies — see below |
| `tier` | 1, 2 or 3 — drives the lede's wording and the chip styling |
| `metric_value` | the phase-appropriate headline figure, or null |
| `metric_key` | `rwm_expected_pct` or `attainment_8_score` — see below |
| `metric_year` | the year the figure is from |
The metric follows the template the page is rendering, not the neighbour's own
phase, so a row of cards never mixes two scales. Primary and **all-through**
pages use `rwm_expected_pct`, matching `PrimarySchoolSections`, which is the
template all-through schools render with; secondary pages use
`attainment_8_score`. Where the neighbour has no value for that key, the card
reads "Not published" rather than falling back to the other key.
`tier` is carried explicitly rather than inferred from the contents of
`shared`, because the frontend needs it for two separate decisions — whether the
lede may claim a similar intake, and whether a chip renders as a brand-tinted
fill or a muted outline — and inferring it from chip count would couple those
decisions to the copy.
Up to six rows of roughly 130 bytes each. It rides in the existing detail payload
rather than a new endpoint because the page already makes exactly one server
fetch for its data, and `/school/[slug]` regenerates at most weekly
(`revalidate = 604800`), so the per-request cost is paid once per school per
week.
**The key is absent, not null, on a backend that does not have this code.** The
frontend treats absent and empty identically, which is what allowed
`NearbyPlaces` to ship without a lockstep deploy of the two images.
`shared` is computed on the backend beside the tier that produced it, not
re-derived on the frontend. Deriving it twice is how a card comes to claim a
match the selection did not actually make.
## Frontend
### Components
`components/school/SimilarSchoolsSection.tsx` — a server component wrapped in
the shared `Section` shell from `sectionShared.tsx`. It renders the heading,
the lede, the card grid, the footer CTA and one caption line. Every
card's title is an `<a>` to the school's canonical slug URL via `schoolUrl()`.
`components/school/AddToCompareButton.tsx` — calls `addSchool` from
`ComparisonProvider` and reports the selection with a `from: 'similar_schools'`
attribution, mirroring `addSchoolFromSearch` in `HomeView.tsx:442`.
`components/school/SimilarSchoolsCarousel.tsx` — the scroller and its arrows. It
takes the server-rendered cards as `children` and the server-rendered heading and
lede as a `header` prop, so those stay server components while the client
component owns only the ref, the scroll handler and the arrows' disabled state.
The split matters: the links — the part with SEO value and the part that must
work without JavaScript — are server-rendered into the initial HTML, and only
the basket interaction and the arrows are hydrated.
### The carousel
**Every card is in the initial HTML.** The arrows scroll a list; they never swap
a view. Six `<a>` elements are in the markup whether or not anything is
hydrated, which is the whole reason the section exists — a paginated widget that
mounts cards on click would put four of the six links beyond a crawler and
beyond a reader with no JavaScript.
So the scroller is a plain overflowing `<ul>` with `scroll-snap-type: x
mandatory`, and the arrows call `scrollBy` on it. With no JavaScript it
degrades to a horizontally scrollable row that still works by touch and by
trackpad. Three cards are visible at desktop width and two below 820px.
**Arrows appear only when there is somewhere to go** — that is, only when more
than three schools were found. Each disables itself at its own end of the
travel.
#### Below 640px the arrows go away
This follows [MOBILE.md](../../../MOBILE.md), which makes 360px the design
floor and mobile the primary target at ≥55% of traffic.
Kept in the heading's flex row at 360px, the two arrow buttons take 96px from a
328px card and crush the lede into a four-line column — measured, not guessed.
And swiping already does what they do. So below 640px the header becomes a
single column, the arrows are not rendered, one card shows at 86% width so the
next one peeks, and the affordance is carried by the right-edge scroll-fade that
MOBILE.md documents for exactly this case:
```css
mask-image: linear-gradient(to right, #000 calc(100% - 28px), transparent);
```
The fade lifts at the end of the travel, where there is nothing left to hint
at. That means the at-end state must be computed whether or not an arrow exists
to consume it — on mobile it drives the mask alone.
**Every interactive element clears 44×44px**, per MOBILE.md's iOS HIG check: the
arrow buttons and the add-to-compare button are both 44px, up from the 40px they
were first drawn at. A card title's own box is shorter than that, but its hit
area is the whole card through the `::after` overlay, so it passes on the target
that actually receives the tap.
**The edge test needs a tolerance, and this is not fussiness.** The scroller
carries 2px of padding so focus rings are not clipped, and scroll-snap treats
that padding as the first card's snap position: a scroller sitting at its start
reports `scrollLeft` of 2, not 0. Sub-pixel rounding moves it again at other
zoom levels. Testing `scrollLeft === 0` therefore leaves the back arrow live and
pointing nowhere on first paint — confirmed in the mockup before it was fixed.
Both ends compare against an 8px tolerance.
**Selecting a school must not move the row.** Adding to the basket re-renders
the footer; the scroll offset lives in the DOM rather than in React state, so
the carousel must not remount or reset on that render. A reader who ticks the
fifth school and is thrown back to the first has been punished for using the
feature.
### Placement and navigation
Rendered as the last section **inside** `SchoolDetailShell`, from both
`PrimarySchoolSections` and `SecondarySchoolSections`. Inside, not after, because
the sticky nav's scroll-spy locates sections with `document.getElementById` and
can only reach a section that lives in the shell.
`NearbyPlaces` stays where it is, outside the shell, immediately below. The
resulting order — this school, then similar schools, then the places containing
them — narrows before it widens, which is the order a reader leaves a page in.
`buildNavItems` and `buildSecondaryNavItems` both gain
`{ id: 'similar', label: 'Similar schools' }`, gated on the section rendering.
The id must match the `Section` id or the scroll-spy silently breaks.
### The comparison CTA
A plain `<a href="/compare?urns=…">`, built from this school's URN plus the
selected ones. `/compare` already parses `urns` from the query string
(`app/(frontend)/compare/page.tsx:55`), so this needs no new compare plumbing.
With nothing selected the CTA is disabled; the button also adds to the shared
basket so the site-wide comparison state stays consistent with what the page
shows.
## Copy, and what the section is allowed to claim
**The lede tracks the deepest tier shown.** At tiers 1–2 it reads "Other primary
schools near X, with a similar intake." Where any card came from tier 3 it drops
"with a similar intake", because for at least one of the cards that is not what
was matched. Six cards make this more likely to fire than three did, which is
correct: a wider net is exactly when the claim needs dropping.
**Chips state only what is shared.** A tier-2 card carries fewer chips rather
than a chip it has not earned; a tier-3 card falls back to the plain phase name,
styled as a muted outline rather than a brand-tinted fill so the difference is
visible at a glance.
**The neighbour's metric carries no valence colour.** Green and terracotta are
reserved site-wide for comparison against the England average. Colouring a
neighbour's figure against this school's would read as ranking the neighbours
against each other, which is precisely the endorsement this section must not
make. The figure sits in neutral ink above a plain "72% at this school"
reference line, and the reader draws their own conclusion.
**A missing figure reads "Not published".** Never 0, never blank, never an
em dash. This follows the same rule the rest of the detail page uses: a school
with no published result has not scored zero.
**There is no "how these are chosen" disclosure.** The method is visible in what
the section already shows — the phase in the lede, the shared characteristics on
each card, the distance above each name — and a collapsed panel restating it
earns less than the space it costs.
**One caption line survives, and only one:** that distances are straight-line
from the school and not road distance. This is not a method note. A reader who
sees "0.6 miles away" and takes it for the walk has been misled by us, and no
other element on the card corrects that. The remaining notes — that listing is
not a recommendation, that special schools only meet special schools — are
statements the selection rules already keep true without being narrated.
## Degradation
| Condition | Behaviour |
|---|---|
| `similar_schools` absent (older backend image) | no section, no nav item |
| fewer than 2 qualifying schools | no section, no nav item |
| this school has no coordinates | no section |
| the helper raises | returns `[]`; the page renders without the section |
The helper is wrapped so a failure inside it never 500s a page that is otherwise
complete — the posture `get_supplementary_data` already takes for its own
queries.
## Testing
**Backend**, in a new `backend/tests/test_similar_schools.py`, against a
synthetic frame rather than live marts:
- a selective school never returns a non-selective one, and vice versa
- a special school returns only special schools; a mainstream school returns none
- a Boys school never returns a Girls school; Mixed matches both
- closed schools and schools without coordinates are never returned
- tier relaxation fills in order, and a school taken at tier 1 is not repeated
- tiers stop relaxing once three are found: four tier-1 matches never open tier 2
- more than six qualifying schools returns the six nearest
- an all-through school is offered on both phase sides
- fewer than two qualifying schools returns `[]`
- distances match a hand-computed haversine for a known pair
**Frontend**, in `nextjs-app/__tests__`:
- the section renders nothing for absent, empty and single-row inputs
- the lede drops "with a similar intake" when any card is tier 3
- a null metric renders "Not published"
- the nav item appears only alongside the section
- every card is in the DOM, including the ones scrolled out of view
- arrows render only when more than three schools were found
jsdom has no layout, so `scrollWidth` and `clientWidth` are both 0 there and the
arrows' disabled state cannot be meaningfully asserted in Jest. That behaviour is
covered in the journey instead, against a real engine, rather than by a unit test
that would pass on a measurement that does not exist.
**E2E**, added to the existing journeys in `e2e/tests` in the same PR, per the
repository's rule on user-facing behaviour:
- the section renders on a known staging URN, with resolving links
- where arrows are present, the back arrow starts disabled and the forward arrow
moves the row
- selecting a school does not reset the scroll position
- add-to-compare reaches `/compare` with the expected `urns`
- at 360, 390 and 430px: no horizontal overflow, every interactive element in the
section clears 44×44px, and no arrows are rendered
The E2E gate runs after merge on this project, so these journeys are not
provable in the PR checks; the PR is verified on the unit tests, and the
journeys are confirmed on the post-merge staging run.
## Out of scope
- A map of the nearby schools. The section is a list; the page already has a map.
- Autoplay, dots, or an infinite loop on the carousel. It is a short list a
reader scans deliberately, not a banner competing for attention, and a row
that moves on its own is a row that moves while someone is reading it.
- Statistical neighbours on deprivation, size or cohort profile. If the tiers
prove too coarse, that is the trigger to move this computation into a dbt mart
— `_similar_schools_payload` is a deliberate seam for exactly that swap.
- Precomputing neighbours in `marts.*`. Rejected for now: a new mart is inert
until Airflow runs, so the feature would ship dark, and every tuning change to
the tiers would become a pipeline round-trip instead of a deploy.
- Any change to `/api/compare`, the compare page, or the comparison basket.