PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m14s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 22s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m24s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 11s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 1m18s
W2 shipped ~5,000 place pages and nothing linked into them. The
location layer pointed down at school pages; school pages pointed
nowhere on the site. Their only anchor was the school's own website, so
the ~27k pages carrying most of the site's inbound authority passed it
straight off-site, and the new corpus was reachable mainly through the
sitemap.
Three things close the loop.
A reverse index over the place registry, places_for_urn, answers which
published places contain a school. Derived from the registry rather
than stored beside it, so the two cannot disagree about which places
exist: a place below the publish threshold is absent from the registry
and therefore never offered as a link. A test asserts that invariant
across every place in a built registry.
GET /api/schools/{urn} gains a `places` array carrying the name, count
and canonical path for each. It rides on the request the page already
makes, so the school page costs no extra round trip. The frontend types
it optional and defaults it to empty, because the two images deploy
separately and a frontend ahead of the API must render without it.
The page gains a "More schools near here" module and a BreadcrumbList.
The module orders narrowest first, because a reader on a school page
wants its town before its county, while the API orders widest first for
the trail. Anchors state their destination's size — "37 schools in
Brentwood" — which is worth more to a reader and a crawler than "see
more". With no published places it renders nothing rather than an empty
heading.
The trail is rooted at the homepage, not /schools. There is no /schools
index page; the location layer lives only at /schools/[place],
/schools/authority/[la] and /schools/near/[outcode]. Rooting it at the
bare path would have opened every breadcrumb with a link to a 404.
Outcodes are omitted from the trail: "schools near CM15" is a real
query and a useful link, but nobody navigates Essex to CM15 to a
school, and a breadcrumb claiming that describes a hierarchy the site
does not have.
School pages also now declare the School type rather than
EducationalOrganization, the parent type that covers universities and
nurseries alike.
The e2e journey asserts the round trip in both directions, following a
place page's own first school so the pair is genuinely related rather
than hardcoded. A one-way link is what already existed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DXnXQKnPpZBBP61fBQiFkq
129 lines
4.9 KiB
Python
129 lines
4.9 KiB
Python
"""Regression tests for GET /api/schools/{urn}.
|
|
|
|
Schools with no performance rows (special post-16 institutions, sixth-form
|
|
centres, PRUs, brand-new schools) come back from the marts LEFT JOIN with
|
|
NaN in every numeric column. The endpoint must still serialize them — a NaN
|
|
that reaches Starlette's JSONResponse raises ValueError (allow_nan=False)
|
|
and the route 500s, which the frontend then renders as a 404.
|
|
"""
|
|
|
|
import numpy as np
|
|
import pandas as pd
|
|
import pytest
|
|
from fastapi.testclient import TestClient
|
|
|
|
|
|
def _no_results_school_df() -> pd.DataFrame:
|
|
"""One school row as produced by the marts query for a school with no
|
|
performance data: GIAS/location fields partly populated, every
|
|
results-linked column NaN (including year)."""
|
|
return pd.DataFrame(
|
|
[
|
|
{
|
|
"urn": 150275,
|
|
"school_name": "West London Performing Arts Academy",
|
|
"phase": "Secondary",
|
|
"school_type": "Special post 16 institution",
|
|
"trust_name": None,
|
|
"religious_denomination": "Does not apply",
|
|
"gender": None,
|
|
"age_range": "16-25",
|
|
"admissions_policy": None,
|
|
"capacity": np.nan,
|
|
"gias_total_pupils": np.nan,
|
|
"headteacher_name": None,
|
|
"website": None,
|
|
"ofsted_grade": np.nan,
|
|
"local_authority": "Ealing",
|
|
"address": "268 Northfield Avenue, London, W5 4UB",
|
|
"postcode": "W5 4UB",
|
|
"latitude": 51.4986,
|
|
"longitude": -0.3148,
|
|
"year": np.nan,
|
|
"total_pupils": np.nan,
|
|
"eligible_pupils": np.nan,
|
|
"rwm_expected_pct": np.nan,
|
|
}
|
|
]
|
|
)
|
|
|
|
|
|
@pytest.fixture()
|
|
def client(monkeypatch):
|
|
from backend import app as app_module
|
|
|
|
monkeypatch.setattr(app_module, "load_school_data", _no_results_school_df)
|
|
monkeypatch.setattr(
|
|
app_module, "get_supplementary_data", lambda db, urn: {}
|
|
)
|
|
return TestClient(app_module.app, raise_server_exceptions=False)
|
|
|
|
|
|
def test_school_without_performance_rows_returns_200(client):
|
|
resp = client.get("/api/schools/150275")
|
|
assert resp.status_code == 200, resp.text
|
|
|
|
|
|
def test_nan_gias_fields_serialize_as_null(client):
|
|
info = client.get("/api/schools/150275").json()["school_info"]
|
|
assert info["capacity"] is None
|
|
assert info["total_pupils"] is None
|
|
assert info["school_name"] == "West London Performing Arts Academy"
|
|
|
|
|
|
# ── Links out to the location layer ─────────────────────────────────────────
|
|
#
|
|
# School pages carried no link into the site at all: the only anchor on the
|
|
# template pointed at the school's own website, so ~27k pages received
|
|
# whatever authority the site had and sent it off-site. `places` is what the
|
|
# link module and the breadcrumb are built from.
|
|
|
|
def test_places_is_present_even_when_the_school_belongs_to_none(client):
|
|
# This fixture's single school cannot clear any publish threshold, so the
|
|
# honest answer is an empty list. The key must still be there: a missing
|
|
# key and "no places" are different things to the page rendering it.
|
|
body = client.get("/api/schools/150275").json()
|
|
assert body["places"] == []
|
|
|
|
|
|
def test_places_names_only_pages_that_exist(monkeypatch):
|
|
from backend import app as app_module
|
|
from backend.places import MIN_SCHOOLS
|
|
|
|
def _df():
|
|
return pd.DataFrame([
|
|
{
|
|
"urn": 100000 + i,
|
|
"school_name": f"Brentwood School {i}",
|
|
"town": "Brentwood",
|
|
"local_authority": "Essex",
|
|
"postcode": "CM15 8AA",
|
|
"phase": "Primary",
|
|
"year": 202425,
|
|
"rwm_expected_pct": 60.0,
|
|
"attainment_8_score": np.nan,
|
|
"ofsted_grade": 2.0,
|
|
"ofsted_date": None,
|
|
}
|
|
for i in range(MIN_SCHOOLS)
|
|
])
|
|
|
|
monkeypatch.setattr(app_module, "load_school_data", _df)
|
|
monkeypatch.setattr(app_module, "get_supplementary_data", lambda db, urn: {})
|
|
monkeypatch.setattr(app_module, "_place_registry", None)
|
|
client = TestClient(app_module.app, raise_server_exceptions=False)
|
|
|
|
places = client.get("/api/schools/100000").json()["places"]
|
|
assert places, "a school in a published town must offer links"
|
|
|
|
by_kind = {p["kind"]: p for p in places}
|
|
assert by_kind["town"]["url"] == "/schools/brentwood"
|
|
assert by_kind["authority"]["url"] == "/schools/authority/essex"
|
|
|
|
# Every entry carries what the link text needs, and a count, so the anchor
|
|
# can say what it leads to rather than "click here".
|
|
for place in places:
|
|
assert place["name"]
|
|
assert place["count"] >= 1
|
|
assert place["url"].startswith("/schools/")
|