feat(seo): link school pages into the location layer
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m14s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 22s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m24s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 11s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 1m18s
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m14s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 22s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m24s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 11s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 1m18s
W2 shipped ~5,000 place pages and nothing linked into them. The
location layer pointed down at school pages; school pages pointed
nowhere on the site. Their only anchor was the school's own website, so
the ~27k pages carrying most of the site's inbound authority passed it
straight off-site, and the new corpus was reachable mainly through the
sitemap.
Three things close the loop.
A reverse index over the place registry, places_for_urn, answers which
published places contain a school. Derived from the registry rather
than stored beside it, so the two cannot disagree about which places
exist: a place below the publish threshold is absent from the registry
and therefore never offered as a link. A test asserts that invariant
across every place in a built registry.
GET /api/schools/{urn} gains a `places` array carrying the name, count
and canonical path for each. It rides on the request the page already
makes, so the school page costs no extra round trip. The frontend types
it optional and defaults it to empty, because the two images deploy
separately and a frontend ahead of the API must render without it.
The page gains a "More schools near here" module and a BreadcrumbList.
The module orders narrowest first, because a reader on a school page
wants its town before its county, while the API orders widest first for
the trail. Anchors state their destination's size — "37 schools in
Brentwood" — which is worth more to a reader and a crawler than "see
more". With no published places it renders nothing rather than an empty
heading.
The trail is rooted at the homepage, not /schools. There is no /schools
index page; the location layer lives only at /schools/[place],
/schools/authority/[la] and /schools/near/[outcode]. Rooting it at the
bare path would have opened every breadcrumb with a link to a 404.
Outcodes are omitted from the trail: "schools near CM15" is a real
query and a useful link, but nobody navigates Essex to CM15 to a
school, and a breadcrumb claiming that describes a hierarchy the site
does not have.
School pages also now declare the School type rather than
EducationalOrganization, the parent type that covers universities and
nurseries alike.
The e2e journey asserts the round trip in both directions, following a
place page's own first school so the pair is genuinely related rather
than hardcoded. A one-way link is what already existed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DXnXQKnPpZBBP61fBQiFkq
This commit is contained in:
1 parent
47f3591ed8
commit
7f5f0fb676
12 files changed
+522
-4
No files matched your search
@@ -8,7 +8,7 @@ import numpy as np
|
||||
import pandas as pd
|
||||
import pytest
|
||||
|
||||
from backend.places import MIN_SCHOOLS, build_place_registry
|
||||
from backend.places import MIN_SCHOOLS, build_place_registry, places_for_urn
|
||||
|
||||
|
||||
def _df(rows: list[dict]) -> pd.DataFrame:
|
||||
@@ -418,3 +418,54 @@ def test_an_authority_still_publishes_phase_variants():
|
||||
and /schools/authority/[la]/[phase] is the route that serves it."""
|
||||
reg = build_place_registry(_df(_town(MIN_SCHOOLS, "Maidstone", "Kent")))
|
||||
assert reg["authority:kent"].publishes_phase("primary")
|
||||
|
||||
|
||||
# ── The reverse index: which published places contain a school ──────────────
|
||||
#
|
||||
# School pages link out to the location layer through this. It is the whole
|
||||
# point of the index: before it, ~27k school pages linked to nothing on the
|
||||
# site and stranded whatever authority they held.
|
||||
|
||||
def test_a_school_resolves_to_every_published_place_containing_it():
|
||||
reg = build_place_registry(_df(_town(MIN_SCHOOLS, "Brentwood", "Essex")))
|
||||
places = places_for_urn(reg, 100000)
|
||||
|
||||
kinds = {p.kind for p in places}
|
||||
assert "town" in kinds
|
||||
assert "authority" in kinds
|
||||
|
||||
|
||||
def test_a_school_in_an_unpublished_town_still_resolves_to_its_authority():
|
||||
# A town below the threshold has no page, so there is no link to offer —
|
||||
# but the authority above it clears the threshold on the same schools and
|
||||
# is where that reader should be sent.
|
||||
reg = build_place_registry(_df(
|
||||
_town(MIN_SCHOOLS - 1, "Tinytown", "Essex")
|
||||
+ _town(MIN_SCHOOLS, "Brentwood", "Essex", start=200000)
|
||||
))
|
||||
places = places_for_urn(reg, 100000)
|
||||
|
||||
# The town is below the threshold, so it has no page and must not be
|
||||
# offered as a link. The authority above it does, and is the right target.
|
||||
assert all(p.slug != "tinytown" for p in places)
|
||||
assert "authority" in {p.kind for p in places}
|
||||
|
||||
|
||||
def test_an_unknown_urn_resolves_to_nothing_rather_than_raising():
|
||||
# A school page renders for any URN the API knows; the link module is not
|
||||
# entitled to take the page down when it has nothing to say.
|
||||
reg = build_place_registry(_df(_town(MIN_SCHOOLS, "Brentwood", "Essex")))
|
||||
assert places_for_urn(reg, 999999) == ()
|
||||
|
||||
|
||||
def test_the_index_is_consistent_with_the_registry_it_was_built_from():
|
||||
# The invariant that matters: a link module must never offer a place whose
|
||||
# page does not exist, and never omit one that does.
|
||||
reg = build_place_registry(_df(
|
||||
_town(MIN_SCHOOLS, "Brentwood", "Essex")
|
||||
+ _town(MIN_SCHOOLS, "Bedford", "Bedford", start=300000)
|
||||
))
|
||||
for key, place in reg.items():
|
||||
for urn in place.urns:
|
||||
assert place in places_for_urn(reg, urn), (
|
||||
f"{urn} is in {key} but the index does not say so")
|
||||
@@ -69,3 +69,60 @@ def test_nan_gias_fields_serialize_as_null(client):
|
||||
assert info["capacity"] is None
|
||||
assert info["total_pupils"] is None
|
||||
assert info["school_name"] == "West London Performing Arts Academy"
|
||||
|
||||
|
||||
# ── Links out to the location layer ─────────────────────────────────────────
|
||||
#
|
||||
# School pages carried no link into the site at all: the only anchor on the
|
||||
# template pointed at the school's own website, so ~27k pages received
|
||||
# whatever authority the site had and sent it off-site. `places` is what the
|
||||
# link module and the breadcrumb are built from.
|
||||
|
||||
def test_places_is_present_even_when_the_school_belongs_to_none(client):
|
||||
# This fixture's single school cannot clear any publish threshold, so the
|
||||
# honest answer is an empty list. The key must still be there: a missing
|
||||
# key and "no places" are different things to the page rendering it.
|
||||
body = client.get("/api/schools/150275").json()
|
||||
assert body["places"] == []
|
||||
|
||||
|
||||
def test_places_names_only_pages_that_exist(monkeypatch):
|
||||
from backend import app as app_module
|
||||
from backend.places import MIN_SCHOOLS
|
||||
|
||||
def _df():
|
||||
return pd.DataFrame([
|
||||
{
|
||||
"urn": 100000 + i,
|
||||
"school_name": f"Brentwood School {i}",
|
||||
"town": "Brentwood",
|
||||
"local_authority": "Essex",
|
||||
"postcode": "CM15 8AA",
|
||||
"phase": "Primary",
|
||||
"year": 202425,
|
||||
"rwm_expected_pct": 60.0,
|
||||
"attainment_8_score": np.nan,
|
||||
"ofsted_grade": 2.0,
|
||||
"ofsted_date": None,
|
||||
}
|
||||
for i in range(MIN_SCHOOLS)
|
||||
])
|
||||
|
||||
monkeypatch.setattr(app_module, "load_school_data", _df)
|
||||
monkeypatch.setattr(app_module, "get_supplementary_data", lambda db, urn: {})
|
||||
monkeypatch.setattr(app_module, "_place_registry", None)
|
||||
client = TestClient(app_module.app, raise_server_exceptions=False)
|
||||
|
||||
places = client.get("/api/schools/100000").json()["places"]
|
||||
assert places, "a school in a published town must offer links"
|
||||
|
||||
by_kind = {p["kind"]: p for p in places}
|
||||
assert by_kind["town"]["url"] == "/schools/brentwood"
|
||||
assert by_kind["authority"]["url"] == "/schools/authority/essex"
|
||||
|
||||
# Every entry carries what the link text needs, and a count, so the anchor
|
||||
# can say what it leads to rather than "click here".
|
||||
for place in places:
|
||||
assert place["name"]
|
||||
assert place["count"] >= 1
|
||||
assert place["url"].startswith("/schools/")
|
||||
Reference in new issue
Block a user