feat(seo): link school pages into the location layer
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m14s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 22s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m24s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 11s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 1m18s

W2 shipped ~5,000 place pages and nothing linked into them. The
location layer pointed down at school pages; school pages pointed
nowhere on the site. Their only anchor was the school's own website, so
the ~27k pages carrying most of the site's inbound authority passed it
straight off-site, and the new corpus was reachable mainly through the
sitemap.

Three things close the loop.

A reverse index over the place registry, places_for_urn, answers which
published places contain a school. Derived from the registry rather
than stored beside it, so the two cannot disagree about which places
exist: a place below the publish threshold is absent from the registry
and therefore never offered as a link. A test asserts that invariant
across every place in a built registry.

GET /api/schools/{urn} gains a `places` array carrying the name, count
and canonical path for each. It rides on the request the page already
makes, so the school page costs no extra round trip. The frontend types
it optional and defaults it to empty, because the two images deploy
separately and a frontend ahead of the API must render without it.

The page gains a "More schools near here" module and a BreadcrumbList.
The module orders narrowest first, because a reader on a school page
wants its town before its county, while the API orders widest first for
the trail. Anchors state their destination's size — "37 schools in
Brentwood" — which is worth more to a reader and a crawler than "see
more". With no published places it renders nothing rather than an empty
heading.

The trail is rooted at the homepage, not /schools. There is no /schools
index page; the location layer lives only at /schools/[place],
/schools/authority/[la] and /schools/near/[outcode]. Rooting it at the
bare path would have opened every breadcrumb with a link to a 404.
Outcodes are omitted from the trail: "schools near CM15" is a real
query and a useful link, but nobody navigates Essex to CM15 to a
school, and a breadcrumb claiming that describes a hierarchy the site
does not have.

School pages also now declare the School type rather than
EducationalOrganization, the parent type that covers universities and
nurseries alike.

The e2e journey asserts the round trip in both directions, following a
place page's own first school so the pair is genuinely related rather
than hardcoded. A one-way link is what already existed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DXnXQKnPpZBBP61fBQiFkq
This commit is contained in:
TudorandClaude Opus 5 committed 2026-09-14 20:45:50 +01:00
1 parent 47f3591ed8
commit 7f5f0fb676
12 files changed
+522 -4

No files matched your search

+30 -1
View File
@@ -38,7 +38,7 @@ from .data_loader import (
)
from .data_loader import get_data_info as get_db_info
from . import flags
from .places import build_place_registry
from .places import build_place_registry, places_for_urn
from .schemas import METRIC_DEFINITIONS, RANKING_COLUMNS, SCHOOL_COLUMNS
from .utils import clean_for_json, convert_to_native
@@ -211,6 +211,28 @@ def _place_url(place) -> str:
return f"/schools/{place.slug}"
def _places_payload(urn: int) -> list[dict]:
"""The published places containing this school, as the school page needs
them: a name to write in the link, a count so the anchor can say what it
leads to, and the canonical path.
`phase_url` is present only where the place publishes a page for this
school's phase, which is the registry's decision alone — repeating the
threshold rule here is how the page and the sitemap would come to disagree.
"""
payload = []
for place in places_for_urn(get_place_registry(), int(urn)):
entry = {
"kind": place.kind,
"slug": place.slug,
"name": place.name,
"count": len(place.urns),
"url": _place_url(place),
}
payload.append(entry)
return payload
def _place_sitemap_rows(kinds: tuple[str, ...]) -> list[str]:
"""A <url> per place, plus a phase variant wherever that phase clears the
threshold on its own.
@@ -902,6 +924,13 @@ async def get_school_details(request: Request, urn: int):
return {
"school_info": school_info,
# Where this school sits in the location layer, for the page's link
# module and breadcrumb. Derived from the same registry the place
# pages and the sitemap use, so a link is never offered for a page
# that does not exist. Empty is a valid answer: a school whose town
# and authority both fall below the publish threshold has nowhere to
# point, and the page renders without the module.
"places": _places_payload(urn),
"yearly_data": clean_for_json(school_data),
# Supplementary data (null if not yet populated by Kestra)
"ofsted": supplementary.get("ofsted"),