Compare commits

..
Author SHA1 Message Date
TudorandClaude Opus 5 b0d5334e06 perf(places): index the reverse lookup, and isolate the registry in tests
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m11s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 18s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m17s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 11s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 1m47s
Two review findings, both confirmed before fixing.

The client fixture in test_school_details.py patched load_school_data
but not _place_registry, which is a module-level cache. A probe settled
it rather than an argument: poisoning the global with a registry built
from data the fixture never saw, then issuing the fixture's own
request, returned that other dataset's places. So the new
`places == []` assertion was satisfied by a stale registry exactly as
well as by the fixture's own data, and proved nothing. Every other test
that touches place data already reset it; the fixture predates places
existing and was never updated. It resets it now.

places_for_urn walked every place in the registry and did a tuple
membership test against each, on /api/schools/{urn}, the site's
highest-traffic endpoint. It now reads a dict built once per registry.
Measured against a synthetic corpus of 27k schools in 1,650 places:
0.118ms per request becomes 0.0001ms, with the index built once in
21ms. Production carries ~5,000 places, so the scan there is larger
again. The absolute saving per request is small; the point is that it
is repeated on every school page view and costs nothing to remove.

The index is cached against the registry by identity rather than
behind a second flag. Anything that drops _place_registry — every test
that touches place data does — gets a fresh registry object, which no
longer matches what the index was built from, so the index rebuilds
with it. A separate _place_index = None would be one more thing to
forget, and a stale reverse index is precisely the first finding's bug
wearing a different hat.

That invalidation has its own test, and the test was checked by
breaking the identity check: five tests fail without it, so three
existing ones were already relying on it too.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DXnXQKnPpZBBP61fBQiFkq
2026-09-14 21:22:39 +01:00
TudorandClaude Opus 5 d65eb58883 fix(seo): build the phase links the docstring already promised
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m12s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 18s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m18s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 11s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 3m48s
Review caught _places_payload documenting a `phase_url` the function
never returned. The docstring was not stray prose: the approved design
included the phase variant — "Primary schools in Beccles" was one of
its four example links — and it was dropped during implementation
without being mentioned. Deleting the sentence would have closed the
report while losing the feature, so the links are built instead.

These are the pages that most needed them. ~950 phase variants were
once reachable by nothing at all: absent from every sitemap and
unlinked from the place page. "Primary schools in brentwood" is the
query they exist to answer.

Membership is read from the registry's own `phase_urns` rather than
re-derived from the school's phase string. The registry already decides
which phases a place publishes and which schools are listed on each, so
asking it is both shorter and the only way the link cannot disagree
with the page it points at. It also means outcodes need no special
case: they carry empty `phase_urns` by design, because nobody searches
"primary schools in SW11", so they report no phase links on their own.

`phases` is a list rather than a single url. An all-through school is
listed on both the primary and secondary pages, so there is no tie to
break and no reason to invent one. Each entry renders directly after
its own place, so "22 primary schools in Brentwood" reads as part of
Brentwood rather than as an unrelated link further along the row.

The e2e journey now follows a phase link where the town publishes one
and asserts it resolves.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DXnXQKnPpZBBP61fBQiFkq
2026-09-14 20:54:59 +01:00
TudorandClaude Opus 5 7f5f0fb676 feat(seo): link school pages into the location layer
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m14s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 22s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m24s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 11s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 1m18s
W2 shipped ~5,000 place pages and nothing linked into them. The
location layer pointed down at school pages; school pages pointed
nowhere on the site. Their only anchor was the school's own website, so
the ~27k pages carrying most of the site's inbound authority passed it
straight off-site, and the new corpus was reachable mainly through the
sitemap.

Three things close the loop.

A reverse index over the place registry, places_for_urn, answers which
published places contain a school. Derived from the registry rather
than stored beside it, so the two cannot disagree about which places
exist: a place below the publish threshold is absent from the registry
and therefore never offered as a link. A test asserts that invariant
across every place in a built registry.

GET /api/schools/{urn} gains a `places` array carrying the name, count
and canonical path for each. It rides on the request the page already
makes, so the school page costs no extra round trip. The frontend types
it optional and defaults it to empty, because the two images deploy
separately and a frontend ahead of the API must render without it.

The page gains a "More schools near here" module and a BreadcrumbList.
The module orders narrowest first, because a reader on a school page
wants its town before its county, while the API orders widest first for
the trail. Anchors state their destination's size — "37 schools in
Brentwood" — which is worth more to a reader and a crawler than "see
more". With no published places it renders nothing rather than an empty
heading.

The trail is rooted at the homepage, not /schools. There is no /schools
index page; the location layer lives only at /schools/[place],
/schools/authority/[la] and /schools/near/[outcode]. Rooting it at the
bare path would have opened every breadcrumb with a link to a 404.
Outcodes are omitted from the trail: "schools near CM15" is a real
query and a useful link, but nobody navigates Essex to CM15 to a
school, and a breadcrumb claiming that describes a hierarchy the site
does not have.

School pages also now declare the School type rather than
EducationalOrganization, the parent type that covers universities and
nurseries alike.

The e2e journey asserts the round trip in both directions, following a
place page's own first school so the pair is genuinely related rather
than hardcoded. A one-way link is what already existed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DXnXQKnPpZBBP61fBQiFkq
2026-09-14 20:45:50 +01:00
tudor 47f3591ed8 Merge pull request 'feat(flags): put /about and /blog behind flags, dark by default' (#144) from feat/about-blog-flags into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 20s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 1m22s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 0s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 2m10s
Reviewed-on: #144
2026-09-08 16:47:18 +00:00
TudorandClaude Opus 5 6fc7fce948 feat(flags): put /about and /blog behind flags, dark by default
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m15s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 20s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m21s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 12s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 1m39s
Both features ship dark. Neither is reachable in an environment where
its flag is off, and every flag in this system starts off, so a deploy
of this commit makes both disappear until someone turns them on
deliberately.

Two independent flags rather than one, which makes blog-on-about-off a
reachable state. That state is the whole reason the change is larger
than four notFound() calls: the blog leans on the About page for its
author identity. The Person entity is anchored at /about#tudor, and
that URL 404s while about_page is dark, so a post published in that
state would claim an author resolving to nothing. Worse than having no
named author. Both bylines fall back to unlinked text and the
BlogPosting attributes to the publisher instead, so every combination
of the two flags renders something correct.

Gated: /about, /blog, /blog/[slug], the RSS feed, both footer links,
and the matching content-sitemap entries. A sitemap must never
advertise a URL that 404s. With both dark it emits a valid empty
urlset rather than a 404, because robots.txt names it unconditionally.

Not gated: /admin. Posts have to be writable before the blog is worth
switching on, so flagging the panel would make the flag unflippable.

getFlags takes a revalidate rather than always using the 300s
constant. Reading a flag pins the calling route to the lowest
revalidate among its fetches, and the footer links live in the root
layout, so a naive gate there would have dropped every school and
place page from a weekly cache to a 5-minute one. The layout passes
604800, the floor those routes already declare, and the build confirms
all four SSG route families still prerender. The cost is one-way
latency: pages follow a flip in minutes, footer links within a week.

The e2e journeys follow the existing paired shape from the
admission_distance flag: a lit journey and a dark one for each flag,
reading state from whether /about and /blog respond rather than from
/api/flags, which another journey asserts is not publicly reachable.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DXnXQKnPpZBBP61fBQiFkq
2026-09-08 17:12:25 +01:00
tudor eb6d918650 Merge pull request 'fix(cms): regenerate the import map so the Content field renders' (#143) from fix/payload-import-map into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 15s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 1m13s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 2m9s
Reviewed-on: #143
2026-09-02 20:12:17 +00:00
25 changed files with 1147 additions and 42 deletions

No files matched your search

+69 -1
View File
@@ -38,7 +38,7 @@ from .data_loader import (
)
from .data_loader import get_data_info as get_db_info
from . import flags
from .places import build_place_registry
from .places import build_place_index, build_place_registry, places_for_urn
from .schemas import METRIC_DEFINITIONS, RANKING_COLUMNS, SCHOOL_COLUMNS
from .utils import clean_for_json, convert_to_native
@@ -65,6 +65,10 @@ _sitemaps: dict[str, str] | None = None
# Built from the same DataFrame the sitemap uses, so places and sitemap can
# never describe different corpora. Reset by the same admin endpoint.
_place_registry: dict | None = None
# Cached beside the registry, and invalidated by identity against it — see
# get_place_index. Never cleared independently.
_place_index: dict | None = None
_place_index_source: dict | None = None
VALID_PLACE_KINDS = ("town", "locality", "authority", "outcode")
@@ -188,6 +192,24 @@ def get_place_registry() -> dict:
return _place_registry
def get_place_index() -> dict:
"""URN → its published places, cached against the registry it came from.
Invalidation is an identity check rather than a second flag to remember to
clear. Anything that drops `_place_registry` — the tests all do — gets a
fresh registry object here, which no longer matches the one the index was
built from, so the index rebuilds with it. A separate `_place_index = None`
would be one more thing to forget, and a stale reverse index is exactly the
bug that would put links to another dataset's places on a school page.
"""
global _place_index, _place_index_source
registry = get_place_registry()
if _place_index is None or _place_index_source is not registry:
_place_index = build_place_index(registry)
_place_index_source = registry
return _place_index
def _urlset(rows: list[str]) -> str:
return "\n".join([
'<?xml version="1.0" encoding="UTF-8"?>',
@@ -211,6 +233,45 @@ def _place_url(place) -> str:
return f"/schools/{place.slug}"
def _places_payload(urn: int) -> list[dict]:
"""The published places containing this school, as the school page needs
them: a name to write in the link, a count so the anchor can say what it
leads to, and the canonical path.
`phases` carries the phase variants this school actually appears on, which
is usually one and is two for an all-through school — it is listed on both
pages, so there is no tie to break.
Membership is read straight from the registry's own `phase_urns` rather
than re-derived from the school's phase string. The registry is the one
place that decides which phases a place publishes and who is on them;
computing it a second time here is how a page comes to link a school to a
phase page that does not list it, or to a route that does not exist. That
is also why outcodes need no special case: they carry empty `phase_urns`,
so they report no phase links on their own.
"""
payload = []
for place in places_for_urn(get_place_index(), int(urn)):
phases = [
{
"phase": phase,
"count": len(phase_urns),
"url": f"{_place_url(place)}/{phase}",
}
for phase, phase_urns in sorted(place.phase_urns.items())
if int(urn) in phase_urns
]
payload.append({
"kind": place.kind,
"slug": place.slug,
"name": place.name,
"count": len(place.urns),
"url": _place_url(place),
"phases": phases,
})
return payload
def _place_sitemap_rows(kinds: tuple[str, ...]) -> list[str]:
"""A <url> per place, plus a phase variant wherever that phase clears the
threshold on its own.
@@ -902,6 +963,13 @@ async def get_school_details(request: Request, urn: int):
return {
"school_info": school_info,
# Where this school sits in the location layer, for the page's link
# module and breadcrumb. Derived from the same registry the place
# pages and the sitemap use, so a link is never offered for a page
# that does not exist. Empty is a valid answer: a school whose town
# and authority both fall below the publish threshold has nowhere to
# point, and the page renders without the module.
"places": _places_payload(urn),
"yearly_data": clean_for_json(school_data),
# Supplementary data (null if not yet populated by Kestra)
"ofsted": supplementary.get("ofsted"),
+17
View File
@@ -57,6 +57,23 @@ REGISTRY: dict[str, Flag] = {
),
added=date(2026, 8, 26),
),
Flag(
name="about_page",
description=(
"The /about page, its footer link, its sitemap entry, and the "
"named-author byline on every blog post."
),
added=date(2026, 9, 8),
),
Flag(
name="blog",
description=(
"The /blog index, post pages, the RSS feed, their footer link "
"and their sitemap entries. Not /admin: posts must be "
"writable before the blog is readable."
),
added=date(2026, 9, 8),
),
)
}
+42
View File
@@ -296,6 +296,48 @@ def _locality_places(df, publishable: set[int],
return out
# Ordered authority → town/locality → outcode, widest first, because that is
# the order a breadcrumb reads. The link module re-sorts for its own purposes.
_PLACE_ORDER = {"authority": 0, "town": 1, "locality": 2, "outcode": 3}
def build_place_index(registry: dict[str, Place]) -> dict[int, tuple[Place, ...]]:
"""URN → the published places containing it, built once per registry.
The reverse of the registry, and the thing school pages link out through.
Derived from the registry rather than maintained beside it, so the two
cannot disagree about which places exist: a place below the publish
threshold is absent from the registry, so it is absent from here too, and
a link is never offered for a page that does not exist.
Built as an index rather than scanned per call because /api/schools/{urn}
is the site's highest-traffic endpoint. Scanning meant walking every place
and doing a tuple membership test against each — on the order of 10^5
comparisons per request, repeated for every school page view. One pass at
registry-build time replaces all of it with a dict lookup.
"""
grouped: dict[int, list[Place]] = {}
for place in registry.values():
for urn in place.urns:
grouped.setdefault(int(urn), []).append(place)
return {
urn: tuple(sorted(places,
key=lambda p: (_PLACE_ORDER.get(p.kind, 9), p.slug)))
for urn, places in grouped.items()
}
def places_for_urn(index: dict[int, tuple[Place, ...]], urn: int) -> tuple[Place, ...]:
"""The published places containing this school, widest first.
Empty is a real answer, not a failure: a school whose town and authority
both fall below the publish threshold has nowhere to link, and the page
renders without the module.
"""
return index.get(int(urn), ())
def build_place_registry(df) -> dict[str, Place]:
"""Every place the site publishes, keyed by "<kind>:<slug>"."""
if df.empty or "urn" not in df.columns:
+78 -1
View File
@@ -8,7 +8,8 @@ import numpy as np
import pandas as pd
import pytest
from backend.places import MIN_SCHOOLS, build_place_registry
from backend.places import (MIN_SCHOOLS, build_place_index,
build_place_registry, places_for_urn)
def _df(rows: list[dict]) -> pd.DataFrame:
@@ -418,3 +419,79 @@ def test_an_authority_still_publishes_phase_variants():
and /schools/authority/[la]/[phase] is the route that serves it."""
reg = build_place_registry(_df(_town(MIN_SCHOOLS, "Maidstone", "Kent")))
assert reg["authority:kent"].publishes_phase("primary")
# ── The reverse index: which published places contain a school ──────────────
#
# School pages link out to the location layer through this. It is the whole
# point of the index: before it, ~27k school pages linked to nothing on the
# site and stranded whatever authority they held.
def test_a_school_resolves_to_every_published_place_containing_it():
reg = build_place_registry(_df(_town(MIN_SCHOOLS, "Brentwood", "Essex")))
places = places_for_urn(build_place_index(reg), 100000)
kinds = {p.kind for p in places}
assert "town" in kinds
assert "authority" in kinds
def test_a_school_in_an_unpublished_town_still_resolves_to_its_authority():
# A town below the threshold has no page, so there is no link to offer —
# but the authority above it clears the threshold on the same schools and
# is where that reader should be sent.
reg = build_place_registry(_df(
_town(MIN_SCHOOLS - 1, "Tinytown", "Essex")
+ _town(MIN_SCHOOLS, "Brentwood", "Essex", start=200000)
))
places = places_for_urn(build_place_index(reg), 100000)
# The town is below the threshold, so it has no page and must not be
# offered as a link. The authority above it does, and is the right target.
assert all(p.slug != "tinytown" for p in places)
assert "authority" in {p.kind for p in places}
def test_an_unknown_urn_resolves_to_nothing_rather_than_raising():
# A school page renders for any URN the API knows; the link module is not
# entitled to take the page down when it has nothing to say.
reg = build_place_registry(_df(_town(MIN_SCHOOLS, "Brentwood", "Essex")))
assert places_for_urn(build_place_index(reg), 999999) == ()
def test_the_index_is_consistent_with_the_registry_it_was_built_from():
# The invariant that matters: a link module must never offer a place whose
# page does not exist, and never omit one that does.
reg = build_place_registry(_df(
_town(MIN_SCHOOLS, "Brentwood", "Essex")
+ _town(MIN_SCHOOLS, "Bedford", "Bedford", start=300000)
))
index = build_place_index(reg)
for key, place in reg.items():
for urn in place.urns:
assert place in places_for_urn(index, urn), (
f"{urn} is in {key} but the index does not say so")
def test_the_index_holds_no_school_the_registry_does_not():
# The reverse direction of the invariant above. An index entry for a URN
# no published place contains would put a link on a page for a place that
# does not list that school.
reg = build_place_registry(_df(
_town(MIN_SCHOOLS, "Brentwood", "Essex")
+ _town(MIN_SCHOOLS - 1, "Tinytown", "Essex", start=400000)
))
index = build_place_index(reg)
for urn, places in index.items():
for place in places:
assert urn in place.urns
assert place.key in reg
def test_the_index_preserves_the_widest_first_order():
# The breadcrumb reads authority then town, and takes this order as given.
reg = build_place_registry(_df(_town(MIN_SCHOOLS, "Brentwood", "Essex")))
kinds = [p.kind for p in places_for_urn(build_place_index(reg), 100000)]
assert kinds.index("authority") < kinds.index("town")
+176
View File
@@ -56,6 +56,11 @@ def client(monkeypatch):
monkeypatch.setattr(
app_module, "get_supplementary_data", lambda db, urn: {}
)
# The place registry is a module-level cache, so without this the endpoint
# answers from whatever registry an earlier test happened to leave behind
# — and a `places == []` assertion is satisfied by a stale registry just
# as well as by this fixture's own data, which makes it prove nothing.
monkeypatch.setattr(app_module, "_place_registry", None)
return TestClient(app_module.app, raise_server_exceptions=False)
@@ -69,3 +74,174 @@ def test_nan_gias_fields_serialize_as_null(client):
assert info["capacity"] is None
assert info["total_pupils"] is None
assert info["school_name"] == "West London Performing Arts Academy"
# ── Links out to the location layer ─────────────────────────────────────────
#
# School pages carried no link into the site at all: the only anchor on the
# template pointed at the school's own website, so ~27k pages received
# whatever authority the site had and sent it off-site. `places` is what the
# link module and the breadcrumb are built from.
def test_places_is_present_even_when_the_school_belongs_to_none(client):
# This fixture's single school cannot clear any publish threshold, so the
# honest answer is an empty list. The key must still be there: a missing
# key and "no places" are different things to the page rendering it.
body = client.get("/api/schools/150275").json()
assert body["places"] == []
def test_places_names_only_pages_that_exist(monkeypatch):
from backend import app as app_module
from backend.places import MIN_SCHOOLS
def _df():
return pd.DataFrame([
{
"urn": 100000 + i,
"school_name": f"Brentwood School {i}",
"town": "Brentwood",
"local_authority": "Essex",
"postcode": "CM15 8AA",
"phase": "Primary",
"year": 202425,
"rwm_expected_pct": 60.0,
"attainment_8_score": np.nan,
"ofsted_grade": 2.0,
"ofsted_date": None,
}
for i in range(MIN_SCHOOLS)
])
monkeypatch.setattr(app_module, "load_school_data", _df)
monkeypatch.setattr(app_module, "get_supplementary_data", lambda db, urn: {})
monkeypatch.setattr(app_module, "_place_registry", None)
client = TestClient(app_module.app, raise_server_exceptions=False)
places = client.get("/api/schools/100000").json()["places"]
assert places, "a school in a published town must offer links"
by_kind = {p["kind"]: p for p in places}
assert by_kind["town"]["url"] == "/schools/brentwood"
assert by_kind["authority"]["url"] == "/schools/authority/essex"
# Every entry carries what the link text needs, and a count, so the anchor
# can say what it leads to rather than "click here".
for place in places:
assert place["name"]
assert place["count"] >= 1
assert place["url"].startswith("/schools/")
def _brentwood_df(phase: str = "Primary", n: int = None):
from backend.places import MIN_SCHOOLS
n = n if n is not None else MIN_SCHOOLS
return lambda: pd.DataFrame([
{
"urn": 100000 + i,
"school_name": f"Brentwood School {i}",
"town": "Brentwood", "local_authority": "Essex",
"postcode": "CM15 8AA", "phase": phase, "year": 202425,
"rwm_expected_pct": 60.0, "attainment_8_score": 50.0,
"ofsted_grade": 2.0, "ofsted_date": None,
}
for i in range(n)
])
def _places_for(monkeypatch, df_factory, urn: int):
from backend import app as app_module
monkeypatch.setattr(app_module, "load_school_data", df_factory)
monkeypatch.setattr(app_module, "get_supplementary_data", lambda db, urn: {})
monkeypatch.setattr(app_module, "_place_registry", None)
client = TestClient(app_module.app, raise_server_exceptions=False)
return client.get(f"/api/schools/{urn}").json()["places"]
def test_a_place_offers_the_phase_page_this_school_appears_on(monkeypatch):
# "primary schools in brentwood" is the query the phase pages exist for,
# and ~950 of them were once reachable by nothing at all.
places = _places_for(monkeypatch, _brentwood_df("Primary"), 100000)
town = next(p for p in places if p["kind"] == "town")
assert town["phases"], "a primary school in a published primary town has a link"
assert town["phases"][0]["url"] == "/schools/brentwood/primary"
assert town["phases"][0]["count"] >= 1
def test_an_all_through_school_offers_both_phase_pages(monkeypatch):
# It genuinely appears on both, so there is no tie to break.
places = _places_for(monkeypatch, _brentwood_df("All-through"), 100000)
town = next(p for p in places if p["kind"] == "town")
assert {p["phase"] for p in town["phases"]} == {"primary", "secondary"}
def test_outcodes_never_offer_a_phase_page(monkeypatch):
# The registry gives outcodes no phase route — nobody searches "primary
# schools in SW11" — and computing them anyway once put a link to a
# nonexistent route on all 1,720 outcode pages.
places = _places_for(monkeypatch, _brentwood_df("Primary"), 100000)
outcode = next((p for p in places if p["kind"] == "outcode"), None)
if outcode is not None:
assert outcode["phases"] == []
def test_a_school_absent_from_the_phase_page_is_not_linked_to_it(monkeypatch):
# The check is URN membership in the registry's own phase list, not a
# re-derivation of the phase mapping. A secondary school must not be sent
# to a primary phase page that does not list it.
from backend.places import MIN_SCHOOLS
def df():
rows = [
{"urn": 100000 + i, "school_name": f"P{i}", "town": "Brentwood",
"local_authority": "Essex", "postcode": "CM15 8AA",
"phase": "Primary", "year": 202425, "rwm_expected_pct": 60.0,
"attainment_8_score": np.nan, "ofsted_grade": 2.0,
"ofsted_date": None}
for i in range(MIN_SCHOOLS)
]
rows.append({
"urn": 900000, "school_name": "Lone Secondary", "town": "Brentwood",
"local_authority": "Essex", "postcode": "CM15 8AA",
"phase": "Secondary", "year": 202425, "rwm_expected_pct": np.nan,
"attainment_8_score": 50.0, "ofsted_grade": 2.0, "ofsted_date": None,
})
return pd.DataFrame(rows)
places = _places_for(monkeypatch, df, 900000)
town = next(p for p in places if p["kind"] == "town")
# The town publishes a primary page, but this secondary school is not on
# it, and there are too few secondaries for a secondary page.
assert town["phases"] == []
def test_the_place_index_rebuilds_when_the_registry_is_replaced(monkeypatch):
"""The reverse index is cached; a stale one would put another dataset's
places on a school page. Invalidation is an identity check against the
registry rather than a second flag, so this asserts the check works."""
from backend import app as app_module
monkeypatch.setattr(app_module, "_place_registry", None)
monkeypatch.setattr(app_module, "_place_index", None)
monkeypatch.setattr(app_module, "_place_index_source", None)
monkeypatch.setattr(app_module, "load_school_data", _brentwood_df("Primary"))
first = app_module.get_place_index()
assert 100000 in first
# Same registry object, so the index is reused rather than rebuilt.
assert app_module.get_place_index() is first
# Drop the registry the way every test that touches place data does. The
# index must follow it, not survive it.
app_module._place_registry = None
monkeypatch.setattr(app_module, "load_school_data",
_brentwood_df("Primary", n=0))
rebuilt = app_module.get_place_index()
assert rebuilt is not first
assert 100000 not in rebuilt, "the index outlived the registry it came from"
+113 -1
View File
@@ -1935,6 +1935,63 @@ async function firstPlaceOfKind(page: Page, kind: string) {
return hit as { kind: string; slug: string; name: string; count: number };
}
/**
* The round trip. Place pages always linked down to school pages; school
* pages linked nowhere on the site, so the ~27k of them that carry most of
* the inbound authority stranded it — their only anchor pointed at the
* school's own website.
*
* Asserting both directions is the point. A one-way link is what already
* existed and is not what this journey is for.
*/
test('a school page links back into the location layer, and the place page links down', async ({ page }) => {
const town = await firstPlaceOfKind(page, 'town');
// Start from the place page and take its first school, so the pair is
// guaranteed to be genuinely related rather than a hardcoded guess.
await page.goto(`/schools/${town.slug}`);
const schoolHref = await page.locator('a[href^="/school/"]').first()
.getAttribute('href');
expect(schoolHref, 'the town page listed no school to follow').toBeTruthy();
await page.goto(schoolHref!);
// Down: the school page must offer a link back to the town it sits in.
const backToTown = page.locator(`a[href="/schools/${town.slug}"]`);
await expect(backToTown).toHaveCount(1);
await expect(backToTown).toBeVisible();
// The anchor says what it leads to, which is worth more than "see more".
await expect(backToTown).toContainText(town.name, { ignoreCase: true });
await expect(backToTown).toContainText(/\d+ schools?/);
// And the breadcrumb resolves the school into a real hierarchy.
const blocks = await page.locator('script[type="application/ld+json"]')
.allTextContents();
const graph = blocks.join(' ');
expect(graph).toContain('"BreadcrumbList"');
// The narrower type, not the EducationalOrganization parent it used to be.
expect(graph).toContain('"School"');
/*
* The phase variants are the pages this most needs to reach: ~950 of them
* were once reachable by nothing at all, absent from every sitemap and
* unlinked from the place page. Conditional because not every school sits
* in a town that publishes one.
*/
const phaseLink = page.locator(`a[href^="/schools/${town.slug}/"]`).first();
if (await phaseLink.count()) {
const phaseHref = await phaseLink.getAttribute('href');
expect((await page.request.get(phaseHref!)).status()).toBe(200);
await expect(phaseLink).toContainText(/primary|secondary/);
}
// Following it lands on a real page, not a 404.
await backToTown.click();
await page.waitForURL(new RegExp(`/schools/${town.slug}$`));
await expect(page.locator('h1')).toContainText(town.name, { ignoreCase: true });
});
for (const [kind, prefix, article] of [
['town', '/schools/', 'a'],
['authority', '/schools/authority/', 'an'],
@@ -2560,8 +2617,56 @@ test('the destinations section never claims a pupil stayed at this school', asyn
* These journeys assert the load-bearing parts of that — a name, a face, the
* honesty claim, and a resolvable Person entity — rather than exact copy,
* which will be edited.
*
* Both are behind flags (about_page, blog), so each has a lit journey and a
* dark one. Flag state is read from the observable effect rather than from
* /api/flags, which the public proxy denies on purpose — the same approach
* distanceFeatureIsOn() takes above.
*/
async function aboutPageIsOn(page: Page): Promise<boolean> {
return (await page.request.get('/about')).ok();
}
async function blogIsOn(page: Page): Promise<boolean> {
return (await page.request.get('/blog')).ok();
}
test('with the about page off, it is absent rather than empty', async ({ page }) => {
test.skip(await aboutPageIsOn(page), 'the about_page flag is on in this environment');
// Dark means the URL does not exist, not that it renders empty: a 404 is
// what stops a crawler keeping the page in its index.
expect((await page.request.get('/about')).status()).toBe(404);
// A footer link into a 404 is the failure this flag has to avoid.
await page.goto('/');
await expect(page.locator('footer a[href="/about"]')).toHaveCount(0);
// And a sitemap must never advertise a URL that 404s.
const sitemap = await page.request.get('/content-sitemap.xml');
expect(await sitemap.text()).not.toContain('/about');
});
test('with the blog off, it is absent rather than empty', async ({ page }) => {
test.skip(await blogIsOn(page), 'the blog flag is on in this environment');
expect((await page.request.get('/blog')).status()).toBe(404);
expect((await page.request.get('/blog/rss.xml')).status()).toBe(404);
await page.goto('/');
await expect(page.locator('footer a[href="/blog"]')).toHaveCount(0);
const sitemap = await page.request.get('/content-sitemap.xml');
expect(await sitemap.text()).not.toContain('/blog');
// The admin panel is deliberately NOT flagged: posts have to be writable
// before the blog is readable, or there is nothing to turn on.
expect((await page.request.get('/admin')).status()).not.toBe(404);
});
test('the about page names a human author and is reachable from the footer', async ({ page }) => {
test.skip(!(await aboutPageIsOn(page)), 'the about_page flag is off in this environment');
await page.goto('/');
const aboutLink = page.locator('footer a[href="/about"]');
await expect(aboutLink).toBeVisible();
@@ -2585,6 +2690,8 @@ test('the about page names a human author and is reachable from the footer', asy
});
test('the blog lists posts and each one renders with a byline', async ({ page }) => {
test.skip(!(await blogIsOn(page)), 'the blog flag is off in this environment');
await page.goto('/blog');
await expect(page.getByRole('heading', { level: 1 })).toBeVisible();
@@ -2612,8 +2719,13 @@ test('the admin panel is not indexable', async ({ page }) => {
test('the content sitemap lists the about page and is advertised in robots', async ({ page }) => {
const sitemap = await page.request.get('/content-sitemap.xml');
// Served whatever the flags say: robots.txt names it unconditionally, and
// with both dark it is a valid empty urlset rather than a 404.
expect(sitemap.ok()).toBeTruthy();
expect(await sitemap.text()).toContain('/about');
if (await aboutPageIsOn(page)) {
expect(await sitemap.text()).toContain('/about');
}
// The school corpus sitemap is proxied from FastAPI; this one is Next's.
// robots.txt must advertise both or the blog never gets discovered.
+15 -2
View File
@@ -27,14 +27,27 @@ describe('BlogPosting structured data', () => {
it('names the same Person entity the about page declares', () => {
// By @id, not by repeating the person: search engines must resolve every
// post and the about page to one author entity, or the site has several.
const ld = blogPostingJsonLd(post);
const ld = blogPostingJsonLd(post, { namedAuthor: true });
expect(ld['@type']).toBe('BlogPosting');
expect(ld.author['@id']).toBe('https://www.schoolcompare.co.uk/about#tudor');
expect(ld.publisher['@id']).toBe('https://www.schoolcompare.co.uk#organization');
});
it('attributes to the organization when the about page is dark', () => {
/*
* The two flags are independent, so blog-on-about-off is a reachable
* state. The Person entity lives at /about#tudor and that URL 404s while
* the flag is dark, so claiming it would declare an author that resolves
* to nothing — worse for the blog's credibility than having no named
* author at all. Attribute to the publisher instead.
*/
const ld = blogPostingJsonLd(post, { namedAuthor: false });
expect(ld.author['@id']).toBe('https://www.schoolcompare.co.uk#organization');
expect(JSON.stringify(ld)).not.toContain('/about');
});
it('carries a self-referencing canonical url and the publish date', () => {
const ld = blogPostingJsonLd(post);
const ld = blogPostingJsonLd(post, { namedAuthor: true });
expect(ld.url).toBe(
'https://www.schoolcompare.co.uk/blog/what-the-data-cannot-tell-you',
);
@@ -0,0 +1,47 @@
/**
* The footer is the only navigational route to /about and /blog, so it is
* where a dark flag would otherwise leave a link into a 404.
*
* Both props default to false. A caller that forgets to pass them hides the
* links, which is the direction that cannot break a page — the same reasoning
* as backend/flags.py's "every flag defaults to False".
*/
import { render, screen } from '@testing-library/react';
import { Footer } from '@/components/Footer';
describe('footer feature links', () => {
it('links to both when both flags are on', () => {
render(<Footer aboutEnabled blogEnabled />);
expect(screen.getByRole('link', { name: /who's behind this/i }))
.toHaveAttribute('href', '/about');
expect(screen.getByRole('link', { name: /^blog$/i }))
.toHaveAttribute('href', '/blog');
});
it('omits the about link when that flag is dark', () => {
render(<Footer blogEnabled />);
expect(screen.queryByRole('link', { name: /who's behind this/i })).toBeNull();
expect(screen.getByRole('link', { name: /^blog$/i })).toBeInTheDocument();
});
it('omits the blog link when that flag is dark', () => {
render(<Footer aboutEnabled />);
expect(screen.queryByRole('link', { name: /^blog$/i })).toBeNull();
expect(screen.getByRole('link', { name: /who's behind this/i }))
.toBeInTheDocument();
});
it('drops the whole section when both are dark, not an empty heading', () => {
// Shipping dark means the footer renders as it did before the feature
// existed, not as a section with its contents removed.
render(<Footer />);
expect(screen.queryByRole('heading', { name: /^about$/i })).toBeNull();
expect(screen.queryByRole('link', { name: /who's behind this/i })).toBeNull();
expect(screen.queryByRole('link', { name: /^blog$/i })).toBeNull();
});
it('defaults to dark when a caller passes nothing', () => {
render(<Footer />);
expect(screen.queryByRole('link', { name: /who's behind this/i })).toBeNull();
});
});
@@ -0,0 +1,92 @@
/**
* The module that ends the stranding: before it, a school page's only anchor
* pointed at the school's own website, so ~27k pages sent authority off-site
* and none of it reached the location layer.
*/
import { render, screen } from '@testing-library/react';
import { NearbyPlaces } from '@/components/school/NearbyPlaces';
const essex = { kind: 'authority', slug: 'essex', name: 'Essex', count: 480, url: '/schools/authority/essex', phases: [] };
const brentwood = { kind: 'town', slug: 'brentwood', name: 'Brentwood', count: 37, url: '/schools/brentwood', phases: [] };
const cm15 = { kind: 'outcode', slug: 'cm15', name: 'CM15', count: 12, url: '/schools/near/cm15', phases: [] };
describe('NearbyPlaces', () => {
it('links to every place the school belongs to', () => {
render(<NearbyPlaces places={[essex, brentwood, cm15]} />);
expect(screen.getByRole('link', { name: /Brentwood/ }))
.toHaveAttribute('href', '/schools/brentwood');
expect(screen.getByRole('link', { name: /Essex/ }))
.toHaveAttribute('href', '/schools/authority/essex');
expect(screen.getByRole('link', { name: /CM15/ }))
.toHaveAttribute('href', '/schools/near/cm15');
});
it('says how many schools each link leads to', () => {
// An anchor that states its destination's size is worth more to a reader
// and to a crawler than "see more".
render(<NearbyPlaces places={[brentwood]} />);
expect(screen.getByRole('link', { name: /37 schools in Brentwood/ }))
.toBeInTheDocument();
});
it('renders nothing at all when the school has no published places', () => {
// Not an empty heading. A school whose town and authority both fall below
// the threshold has nowhere to point, and the page should look as it did
// before the module existed.
const { container } = render(<NearbyPlaces places={[]} />);
expect(container).toBeEmptyDOMElement();
});
it('puts the narrowest place first, which is the most useful link', () => {
// The API orders widest-first for the breadcrumb; a reader on a school
// page wants its town before its county.
render(<NearbyPlaces places={[essex, brentwood, cm15]} />);
const hrefs = screen.getAllByRole('link').map((a) => a.getAttribute('href'));
expect(hrefs.indexOf('/schools/brentwood'))
.toBeLessThan(hrefs.indexOf('/schools/authority/essex'));
});
it('handles a singular count without saying "1 schools"', () => {
render(<NearbyPlaces places={[{ ...brentwood, count: 1 }]} />);
expect(screen.getByRole('link', { name: /1 school in Brentwood/ }))
.toBeInTheDocument();
});
it('links the phase page the school appears on', () => {
// "primary schools in brentwood" is the query these pages exist for.
render(<NearbyPlaces places={[{
...brentwood,
phases: [{ phase: 'primary', count: 22, url: '/schools/brentwood/primary' }],
}]} />);
expect(screen.getByRole('link', { name: /22 primary schools in Brentwood/ }))
.toHaveAttribute('href', '/schools/brentwood/primary');
});
it('links both phase pages for an all-through school', () => {
render(<NearbyPlaces places={[{
...brentwood,
phases: [
{ phase: 'primary', count: 22, url: '/schools/brentwood/primary' },
{ phase: 'secondary', count: 9, url: '/schools/brentwood/secondary' },
],
}]} />);
expect(screen.getByRole('link', { name: /22 primary schools/ })).toBeInTheDocument();
expect(screen.getByRole('link', { name: /9 secondary schools/ })).toBeInTheDocument();
});
it('keeps a phase link next to the place it belongs to', () => {
// Grouping matters: "22 primary schools in Brentwood" directly after
// "37 schools in Brentwood" reads as one place, not two unrelated links.
render(<NearbyPlaces places={[essex, {
...brentwood,
phases: [{ phase: 'primary', count: 22, url: '/schools/brentwood/primary' }],
}]} />);
const hrefs = screen.getAllByRole('link').map((a) => a.getAttribute('href'));
expect(hrefs.indexOf('/schools/brentwood/primary'))
.toBe(hrefs.indexOf('/schools/brentwood') + 1);
});
});
+25 -1
View File
@@ -1,4 +1,4 @@
import { getFlags } from '@/lib/flags';
import { getFlags, FLAGS_REVALIDATE } from '@/lib/flags';
// jsdom provides no global fetch, so there is nothing for jest.spyOn to attach
// to — assign it and restore the original afterwards. This is the first test
@@ -31,4 +31,28 @@ describe('getFlags', () => {
mockFetch(async () => ({ ok: false, status: 503 }));
await expect(getFlags()).resolves.toEqual({});
});
/*
* Reading a flag pins the calling route's ISR floor: Next uses the LOWEST
* revalidate among a route's fetches for the whole route. That is why the
* revalidate is an argument rather than the constant.
*
* Every SEO route here declares `revalidate = 604800`. A gate that read
* flags at the 300s default would drop the whole school and place corpus
* from a weekly cache to a 5-minute one, which is a large origin-load
* regression to pay for a feature flag.
*/
it('reads at the 300s floor by default', async () => {
mockFetch(async () => ({ ok: true, json: async () => ({}) }));
await getFlags();
expect((global.fetch as jest.Mock).mock.calls[0][1])
.toEqual({ next: { revalidate: FLAGS_REVALIDATE } });
});
it('lets a caller pass its own route floor instead', async () => {
mockFetch(async () => ({ ok: true, json: async () => ({}) }));
await getFlags(604800);
expect((global.fetch as jest.Mock).mock.calls[0][1])
.toEqual({ next: { revalidate: 604800 } });
});
});
@@ -0,0 +1,68 @@
/**
* School pages had no BreadcrumbList and no links into the location layer.
* Both are fixed by the same data — the `places` array the API now returns —
* so they are tested together.
*/
import { schoolBreadcrumbJsonLd } from '@/lib/jsonld';
const essex = { kind: 'authority', slug: 'essex', name: 'Essex', count: 480, url: '/schools/authority/essex', phases: [] };
const brentwood = { kind: 'town', slug: 'brentwood', name: 'Brentwood', count: 37, url: '/schools/brentwood', phases: [] };
const outcode = { kind: 'outcode', slug: 'cm15', name: 'CM15', count: 12, url: '/schools/near/cm15', phases: [] };
describe('school breadcrumbs', () => {
it('reads home to authority to town to school', () => {
const ld = schoolBreadcrumbJsonLd({
name: 'Brentwood School', url: '/school/100000-brentwood-school',
places: [essex, brentwood],
});
expect(ld['@type']).toBe('BreadcrumbList');
expect(ld.itemListElement.map((i) => i.name))
.toEqual(['schoolcompare', 'Essex', 'Brentwood', 'Brentwood School']);
expect(ld.itemListElement.map((i) => i.position)).toEqual([1, 2, 3, 4]);
});
it('skips a level the school has no published place for', () => {
// A school whose town falls below the publish threshold has no town page.
// The trail closes over the gap rather than linking to a 404.
const ld = schoolBreadcrumbJsonLd({
name: 'Lone School', url: '/school/1-lone-school', places: [essex],
});
expect(ld.itemListElement.map((i) => i.name))
.toEqual(['schoolcompare', 'Essex', 'Lone School']);
expect(ld.itemListElement.map((i) => i.position)).toEqual([1, 2, 3]);
});
it('omits outcodes, which are not a place a breadcrumb reads through', () => {
// CM15 is a useful link in the module but nonsense in a trail: nobody
// navigates Essex → CM15 → school.
const ld = schoolBreadcrumbJsonLd({
name: 'Brentwood School', url: '/school/100000-brentwood-school',
places: [essex, brentwood, outcode],
});
expect(JSON.stringify(ld)).not.toContain('cm15');
});
it('still produces a valid trail when the school has no places at all', () => {
const ld = schoolBreadcrumbJsonLd({
name: 'Orphan School', url: '/school/2-orphan-school', places: [],
});
expect(ld.itemListElement.map((i) => i.name)).toEqual(['schoolcompare', 'Orphan School']);
});
it('uses absolute urls, as every other entity on the site does', () => {
const ld = schoolBreadcrumbJsonLd({
name: 'Brentwood School', url: '/school/100000-brentwood-school',
places: [essex, brentwood],
});
for (const item of ld.itemListElement) {
expect(item.item).toMatch(/^https:\/\/www\.schoolcompare\.co\.uk\//);
}
// The root is the homepage: there is no /schools index page to link to.
expect(ld.itemListElement[0].item).toBe('https://www.schoolcompare.co.uk/');
});
});
+14 -1
View File
@@ -1,6 +1,8 @@
import type { Metadata } from 'next';
import Image from 'next/image';
import { notFound } from 'next/navigation';
import { absoluteUrl } from '@/lib/site';
import { getFlags } from '@/lib/flags';
import { personJsonLd, organizationJsonLd } from '@/lib/jsonld';
import styles from './About.module.css';
@@ -11,7 +13,18 @@ export const metadata: Metadata = {
alternates: { canonical: absoluteUrl('/about') },
};
export default function AboutPage() {
/*
* Gated on about_page. The default 300s read is the right floor here: this
* page declares no revalidate of its own, so nothing is lost by it, and a flip
* lands within five minutes.
*
* notFound(), not a redirect: while the flag is dark this URL does not exist,
* and a 404 is what tells a crawler not to keep it.
*/
export default async function AboutPage() {
const flags = await getFlags();
if (flags.about_page !== true) notFound();
const jsonLd = {
'@context': 'https://schema.org',
'@graph': [personJsonLd(), organizationJsonLd()],
+21 -3
View File
@@ -7,6 +7,7 @@ import type { JSXConvertersFunction } from '@payloadcms/richtext-lexical/react';
import { getCachedPayload } from '@/lib/payload';
import type { Post, Media } from '@/payload-types';
import { absoluteUrl } from '@/lib/site';
import { getFlags } from '@/lib/flags';
import {
blogPostingJsonLd,
breadcrumbJsonLd,
@@ -112,6 +113,15 @@ export default async function PostPage(
{ params }: { params: Promise<{ slug: string }> },
) {
const { slug } = await params;
/*
* Flags read at this route's own declared floor, so gating costs it nothing.
* Checked before the post is fetched: a dark blog should not query Payload.
*/
const flags = await getFlags(3600);
if (flags.blog !== true) notFound();
const namedAuthor = flags.about_page === true;
const post = await findPost(slug);
if (!post) notFound();
@@ -120,10 +130,15 @@ export default async function PostPage(
const jsonLd = {
'@context': 'https://schema.org',
/*
* The Person entity is anchored at /about#tudor, so it is declared only
* when that page exists. Claiming an author whose URL 404s is a worse
* signal than attributing the post to the publisher.
*/
'@graph': [
blogPostingJsonLd(summary),
blogPostingJsonLd(summary, { namedAuthor }),
breadcrumbJsonLd(summary),
personJsonLd(),
...(namedAuthor ? [personJsonLd()] : []),
organizationJsonLd(),
],
};
@@ -142,7 +157,10 @@ export default async function PostPage(
<h1 className={styles.heading}>{summary.title}</h1>
<p className={styles.byline}>
By <Link href="/about" className={styles.link}>Tudor</Link>
{/* Unlinked while about_page is dark; the flags are independent. */}
By {namedAuthor
? <Link href="/about" className={styles.link}>Tudor</Link>
: 'Tudor'}
{' · '}
<time dateTime={summary.publishedAt}>
{new Date(summary.publishedAt).toLocaleDateString('en-GB', {
+10 -1
View File
@@ -1,7 +1,9 @@
import type { Metadata } from 'next';
import Link from 'next/link';
import { notFound } from 'next/navigation';
import { getCachedPayload } from '@/lib/payload';
import { absoluteUrl } from '@/lib/site';
import { getFlags } from '@/lib/flags';
import styles from './Blog.module.css';
/*
@@ -31,6 +33,9 @@ function formatDate(value: string) {
}
export default async function BlogIndexPage() {
const flags = await getFlags();
if (flags.blog !== true) notFound();
const payload = await getCachedPayload();
const { docs } = await payload.find({
collection: 'posts',
@@ -48,7 +53,11 @@ export default async function BlogIndexPage() {
<p className={styles.standfirst}>
What school performance data shows, what it doesn&apos;t, and how to
read it without being misled. Written by{' '}
<Link href="/about" className={styles.link}>Tudor</Link>.
{/* Plain text when about_page is dark: the two flags are
independent, so this link would otherwise point at a 404. */}
{flags.about_page === true
? <Link href="/about" className={styles.link}>Tudor</Link>
: 'Tudor'}.
</p>
</header>
@@ -1,5 +1,6 @@
import { getCachedPayload } from '@/lib/payload';
import { absoluteUrl } from '@/lib/site';
import { getFlags } from '@/lib/flags';
/*
* Dynamic, not ISR.
@@ -18,6 +19,11 @@ function escapeXml(value: string): string {
}
export async function GET() {
// A dark blog has no feed. 404 rather than an empty channel: an empty feed
// is a live feed with nothing in it, which a reader would keep polling.
const flags = await getFlags();
if (flags.blog !== true) return new Response('Not found', { status: 404 });
const payload = await getCachedPayload();
const { docs } = await payload.find({
collection: 'posts',
@@ -8,6 +8,7 @@
*/
import { getCachedPayload } from '@/lib/payload';
import { absoluteUrl } from '@/lib/site';
import { getFlags } from '@/lib/flags';
/*
* Dynamic, not ISR.
@@ -21,18 +22,33 @@ import { absoluteUrl } from '@/lib/site';
export const dynamic = 'force-dynamic';
export async function GET() {
const payload = await getCachedPayload();
const { docs } = await payload.find({
collection: 'posts',
where: { _status: { equals: 'published' } },
sort: '-publishedAt',
limit: 500,
depth: 0,
});
/*
* A dark page must not be advertised. Submitting a URL that 404s is the one
* thing a sitemap is not allowed to do, so each entry is gated on the same
* flag that gates the page itself.
*
* With both flags dark this emits a valid, empty <urlset> rather than a 404:
* robots.txt names this sitemap unconditionally, and an empty sitemap is a
* well-formed statement that there is nothing here yet.
*/
const flags = await getFlags();
const aboutEnabled = flags.about_page === true;
const blogEnabled = flags.blog === true;
// Only query Payload when the blog is actually being advertised.
const docs = blogEnabled
? (await (await getCachedPayload()).find({
collection: 'posts',
where: { _status: { equals: 'published' } },
sort: '-publishedAt',
limit: 500,
depth: 0,
})).docs
: [];
const urls: Array<{ loc: string; lastmod: string | null }> = [
{ loc: absoluteUrl('/about'), lastmod: null },
{ loc: absoluteUrl('/blog'), lastmod: null },
...(aboutEnabled ? [{ loc: absoluteUrl('/about'), lastmod: null }] : []),
...(blogEnabled ? [{ loc: absoluteUrl('/blog'), lastmod: null }] : []),
...docs.map((post) => ({
loc: absoluteUrl(`/blog/${post.slug}`),
lastmod: new Date(String(post.updatedAt ?? post.publishedAt)).toISOString(),
+23 -2
View File
@@ -7,6 +7,7 @@ import { ComparisonToast } from '@/components/ComparisonToast';
import { RouteTrail } from '@/components/RouteTrail';
import { ComparisonProvider } from '@/context/ComparisonProvider';
import { SITE_URL } from '@/lib/site';
import { getFlags } from '@/lib/flags';
import './globals.css';
// Manrope carries headings and key messaging — the guideline's "friendly,
@@ -78,11 +79,28 @@ export const metadata: Metadata = {
},
};
export default function RootLayout({
/*
* The footer's About and Blog links are flagged, which makes this the one
* place on the site that reads a flag on every route.
*
* 604800 is deliberate and load-bearing: it is the revalidate every SEO route
* here already declares. Next pins a route to the LOWEST revalidate among its
* fetches, so reading flags at the 300s default would drop the whole school
* and place corpus from a weekly cache to a 5-minute one — a large origin-load
* regression to hide two footer links.
*
* The cost is latency in one direction only. The pages themselves read the
* same flags at their own floors and flip within minutes; the footer links
* follow within a week. Turning a feature on early therefore shows the page
* before its footer link, which is harmless. Turning one off leaves a link to
* a 404 until the cache turns over, so a rollback that matters wants a purge.
*/
export default async function RootLayout({
children,
}: Readonly<{
children: React.ReactNode;
}>) {
const flags = await getFlags(604800);
return (
// The font variable classes must sit on <html>, not <body>. globals.css
// declares --font-display on :root as var(--font-manrope) and --font-ui as
@@ -126,7 +144,10 @@ export default function RootLayout({
{children}
</main>
<ComparisonToast />
<Footer />
<Footer
aboutEnabled={flags.about_page === true}
blogEnabled={flags.blog === true}
/>
</ComparisonProvider>
</body>
</html>
@@ -7,6 +7,8 @@
import { fetchSchoolDetails, fetchSchools, fetchNationalAverages } from '@/lib/api';
import { notFound, redirect } from 'next/navigation';
import { SchoolDetailShell } from '@/components/school/SchoolDetailShell';
import { NearbyPlaces } from '@/components/school/NearbyPlaces';
import { schoolBreadcrumbJsonLd, type SchoolPlace } from '@/lib/jsonld';
import { PrimarySchoolSections } from '@/components/school/PrimarySchoolSections';
import { SecondarySchoolSections } from '@/components/school/SecondarySchoolSections';
import {
@@ -149,6 +151,10 @@ export default async function SchoolPage({ params }: SchoolPageProps) {
}
const { school_info, yearly_data, absence_data, ofsted, census, admissions, admissions_history, admission_distance, deprivation, finance, destinations } = data;
// Absent on an older API build; the module and the trail both degrade to
// nothing rather than throwing, which is how this shipped without a
// lockstep deploy of the two images.
const places: SchoolPlace[] = data.places ?? [];
// Redirect bare URN to canonical slug URL
const canonicalSlug = schoolUrl(urn, school_info.school_name).replace('/school/', '');
@@ -185,10 +191,19 @@ export default async function SchoolPage({ params }: SchoolPageProps) {
const primaryNavItems = buildNavItems(primaryFlags, navInput);
const secondaryNavItems = buildSecondaryNavItems(secondaryFlags, navInput);
// Generate JSON-LD structured data for SEO
/*
* `School`, not `EducationalOrganization`.
*
* Both are valid, but EducationalOrganization is the parent type covering
* universities, training providers and nurseries alike. School is the
* specific one, and a type that says what the page is about is the whole
* point of declaring it. Google's own guidance treats the narrower type as
* the correct choice where it applies.
*/
const structuredData = {
'@context': 'https://schema.org',
'@type': 'EducationalOrganization',
'@graph': [{
'@type': 'School',
name: school_info.school_name,
identifier: school_info.urn.toString(),
...(school_info.address && {
@@ -210,6 +225,15 @@ export default async function SchoolPage({ params }: SchoolPageProps) {
...(school_info.school_type && {
additionalType: school_info.school_type,
}),
},
// The trail the page sits at the end of. School pages carried no
// breadcrumb at all, while every place page already emitted one.
schoolBreadcrumbJsonLd({
name: school_info.school_name,
url: `/school/${slug}`,
places,
}),
],
};
return (
@@ -264,6 +288,7 @@ export default async function SchoolPage({ params }: SchoolPageProps) {
/>
</SchoolDetailShell>
)}
<NearbyPlaces places={places} />
</>
);
}
+31 -12
View File
@@ -10,7 +10,17 @@
import { LogoMark } from './Logo';
import styles from './Footer.module.css';
export function Footer() {
/**
* Both default to false so a caller that forgets a prop hides the link rather
* than pointing it at a page that 404s. Same reasoning as backend/flags.py:
* "Every flag defaults to False."
*/
interface FooterProps {
aboutEnabled?: boolean;
blogEnabled?: boolean;
}
export function Footer({ aboutEnabled = false, blogEnabled = false }: FooterProps = {}) {
const currentYear = new Date().getFullYear();
return (
@@ -94,17 +104,26 @@ export function Footer() {
</ul>
</div>
<div className={styles.section}>
<h4 className={styles.sectionTitle}>About</h4>
<ul className={styles.links}>
{/* The only route to a named human. Deliberately not in the nav:
the mobile bottom bar already carries four items, and both of
these are lower intent than any of them. Post bylines link
here too, which is where a reader actually asks the question. */}
<li><a href="/about" className={styles.link}>Who&apos;s behind this</a></li>
<li><a href="/blog" className={styles.link}>Blog</a></li>
</ul>
</div>
{/* Dropped entirely when both flags are dark, rather than left as an
empty heading: shipping dark means the footer renders as it did
before the feature existed. */}
{(aboutEnabled || blogEnabled) && (
<div className={styles.section}>
<h4 className={styles.sectionTitle}>About</h4>
<ul className={styles.links}>
{/* The only route to a named human. Deliberately not in the nav:
the mobile bottom bar already carries four items, and both of
these are lower intent than any of them. Post bylines link
here too, which is where a reader actually asks the question. */}
{aboutEnabled && (
<li><a href="/about" className={styles.link}>Who&apos;s behind this</a></li>
)}
{blogEnabled && (
<li><a href="/blog" className={styles.link}>Blog</a></li>
)}
</ul>
</div>
)}
</div>
<div className={styles.bottom}>
@@ -0,0 +1,53 @@
/* Tokens only — the same vocabulary schoolSections.module.css uses, so the
module follows both themes without a rule of its own. No hardcoded colour
appears here; darkThemeSafety asserts that across the codebase. */
.section {
margin-top: 2rem;
}
/* Matches .sectionTitle in schoolSections.module.css, including the brand
rule before the text, so this reads as one more section of the page
rather than a footer bolted underneath it. */
.heading {
font-size: 1.125rem;
font-weight: 600;
color: var(--text-primary);
margin-bottom: 0.875rem;
padding-bottom: 0.5rem;
border-bottom: 2px solid var(--border);
font-family: var(--font-display);
display: flex;
align-items: center;
gap: 0.375rem;
}
.heading::before {
content: "";
display: inline-block;
width: 3px;
height: 1em;
background: var(--brand);
border-radius: 2px;
flex-shrink: 0;
}
.list {
display: flex;
flex-wrap: wrap;
gap: 0.5rem 1.25rem;
list-style: none;
margin: 0;
padding: 0;
}
.link {
color: var(--brand-strong);
font-weight: 500;
text-decoration: underline;
text-underline-offset: 2px;
}
.link:hover {
text-decoration-thickness: 2px;
}
@@ -0,0 +1,67 @@
import Link from 'next/link';
import type { SchoolPlace, SchoolPhasePage } from '@/lib/jsonld';
import styles from './NearbyPlaces.module.css';
/**
* Links from a school page into the location layer.
*
* This exists for a structural reason rather than a decorative one. Before
* it, the only anchor on a school page pointed at the school's own website,
* so the ~27k pages that carry most of the site's inbound authority passed it
* straight off-site and none of it reached the place pages. These links are
* what circulate it instead.
*
* Every entry comes from the place registry via the API, so a link is only
* ever offered for a page that exists: a place below the publish threshold is
* absent from the registry and therefore absent here.
*/
/** Narrowest first: a reader on a school page wants its town before its
* county. The API orders widest-first because that is what the breadcrumb
* reads, so the two orders are deliberately different. */
const ORDER: Record<string, number> = {
town: 0, locality: 0, outcode: 1, authority: 2,
};
function label(place: SchoolPlace): string {
const noun = place.count === 1 ? 'school' : 'schools';
const preposition = place.kind === 'outcode' ? 'near' : 'in';
return `${place.count} ${noun} ${preposition} ${place.name}`;
}
/** "22 primary schools in Brentwood" — the phrasing the query itself uses. */
function phaseLabel(place: SchoolPlace, page: SchoolPhasePage): string {
const noun = page.count === 1 ? 'school' : 'schools';
return `${page.count} ${page.phase} ${noun} in ${place.name}`;
}
export function NearbyPlaces({ places }: { places: SchoolPlace[] }) {
if (places.length === 0) return null;
const sorted = [...places].sort(
(a, b) => (ORDER[a.kind] ?? 9) - (ORDER[b.kind] ?? 9),
);
return (
<section className={styles.section} aria-labelledby="nearby-places">
<h2 id="nearby-places" className={styles.heading}>More schools near here</h2>
<ul className={styles.list}>
{sorted.flatMap((place) => [
<li key={`${place.kind}:${place.slug}`}>
<Link href={place.url} className={styles.link}>{label(place)}</Link>
</li>,
/* Immediately after its own place, so "22 primary schools in
Brentwood" reads as part of Brentwood rather than as an
unrelated link further down the row. */
...place.phases.map((page) => (
<li key={`${place.kind}:${place.slug}:${page.phase}`}>
<Link href={page.url} className={styles.link}>
{phaseLabel(place, page)}
</Link>
</li>
)),
])}
</ul>
</section>
);
}
+20
View File
@@ -4,6 +4,26 @@ The blog is Payload CMS, running inside the Next.js app. There is no separate
service and no second deploy. Writing a post is done in the browser and takes
effect on the live site within seconds.
## Before any of this: the flags
`/blog` and `/about` are behind feature flags (`blog` and `about_page`), and
every flag in this system starts off. While `blog` is dark, `/blog`, every post
page and the RSS feed return 404, and neither appears in the sitemap or the
footer. Posts still save normally, because `/admin` is deliberately **not**
flagged: you have to be able to write a post before there is anything worth
switching on.
So a new post published to a dark blog is invisible, and that is working as
intended, not a bug. Flip the flag in Unleash when the content is ready.
Staging and production hold their own values (`development` and `production`
environments), so you can light it on staging first. The page follows a flip
within five minutes; the footer link takes up to a week, because it renders in
the root layout and is cached at the same weekly floor as the school corpus.
Both flags are temporary scaffolding, like every flag here: a test starts
failing once one is older than 90 days, at which point either the feature is
permanent and the flag comes out, or it was never going to ship.
## Signing in
`https://www.schoolcompare.co.uk/admin`, one account, no registration. If you
+13 -3
View File
@@ -26,11 +26,21 @@ export const FLAGS_REVALIDATE = 300;
const API = process.env.FASTAPI_URL || process.env.NEXT_PUBLIC_API_URL
|| 'http://localhost:8000/api';
/** Every flag and its value. Never throws: an unreadable flag is a dark one. */
export async function getFlags(): Promise<Flags> {
/**
* Every flag and its value. Never throws: an unreadable flag is a dark one.
*
* `revalidate` is the caller's, because reading flags pins the whole route to
* the lowest revalidate among its fetches. Every SEO route here declares
* 604800; gating one at the 300s default would drop it from a weekly cache to
* a 5-minute one. Pass the route's own floor and gating costs it nothing.
*
* The trade is flag-flip latency: a route that revalidates weekly takes up to
* a week to notice a flip. Pass a smaller number where a flip must land fast.
*/
export async function getFlags(revalidate: number = FLAGS_REVALIDATE): Promise<Flags> {
try {
const res = await fetch(`${API}/flags`, {
next: { revalidate: FLAGS_REVALIDATE },
next: { revalidate },
});
if (!res.ok) return {};
return await res.json();
+83 -2
View File
@@ -44,19 +44,100 @@ interface PostSummary {
* References the Person and Organization by @id rather than repeating them, so
* search engines resolve every post and the About page to the one author
* entity. Repeating the shape would declare several people with one name.
*
* `namedAuthor` is the about_page flag. The Person entity is anchored at
* /about#tudor, and that URL 404s while the flag is dark, so a post published
* in that state must not claim it: an author @id resolving to nothing is a
* worse signal than no named author. It falls back to the publisher, which is
* always live. The parameter is required rather than defaulted because every
* call site has the flag to hand and the wrong default is silent.
*/
export function blogPostingJsonLd(post: PostSummary) {
export function blogPostingJsonLd(
post: PostSummary,
{ namedAuthor }: { namedAuthor: boolean },
) {
return {
'@type': 'BlogPosting',
headline: post.title,
description: post.excerpt,
url: absoluteUrl(`/blog/${post.slug}`),
datePublished: post.publishedAt,
author: { '@id': `${SITE_URL}/about#tudor` },
author: {
'@id': namedAuthor ? `${SITE_URL}/about#tudor` : `${SITE_URL}#organization`,
},
publisher: { '@id': `${SITE_URL}#organization` },
} as const;
}
/**
* A place the location layer publishes a page for, as the school API reports
* it. `count` is what lets a link say "All 37 schools in Brentwood" rather
* than "click here".
*/
export interface SchoolPhasePage {
phase: string;
count: number;
url: string;
}
export interface SchoolPlace {
kind: string;
slug: string;
name: string;
count: number;
url: string;
/**
* The phase variants this school is actually listed on: usually one, two
* for an all-through school, none for an outcode, which publishes no phase
* route. Decided by the place registry, never re-derived here.
*/
phases: SchoolPhasePage[];
}
/**
* The trail a school page sits at the end of: Schools → authority → town.
*
* Only authority and town/locality appear. An outcode is a useful link in the
* module beside this — a parent does search "schools near CM15" — but it is
* not a step anyone navigates through, and a breadcrumb that claims otherwise
* describes a hierarchy the site does not have.
*
* Levels are skipped rather than faked. A school whose town falls below the
* publish threshold has no town page, so the trail closes over the gap; the
* alternative is a breadcrumb linking to a 404.
*/
export function schoolBreadcrumbJsonLd(
school: { name: string; url: string; places: SchoolPlace[] },
) {
/*
* Rooted at the homepage, not at /schools. There is no /schools index page
* — the location layer is /schools/[place], /schools/authority/[la] and
* /schools/near/[outcode], with nothing at the bare path — so a trail
* starting there would open with a link to a 404.
*/
const trail: Array<{ name: string; url: string }> = [
{ name: 'schoolcompare', url: '/' },
];
const authority = school.places.find((p) => p.kind === 'authority');
if (authority) trail.push({ name: authority.name, url: authority.url });
const town = school.places.find((p) => p.kind === 'town' || p.kind === 'locality');
if (town) trail.push({ name: town.name, url: town.url });
trail.push({ name: school.name, url: school.url });
return {
'@type': 'BreadcrumbList',
itemListElement: trail.map((step, index) => ({
'@type': 'ListItem',
position: index + 1,
name: step.name,
item: absoluteUrl(step.url),
})),
} as const;
}
export function breadcrumbJsonLd(post: PostSummary) {
return {
'@type': 'BreadcrumbList',
+11
View File
@@ -1,3 +1,5 @@
import type { SchoolPlace } from '@/lib/jsonld';
/**
* TypeScript type definitions for SchoolCompare API
* Generated from backend/models.py and backend/schemas.py
@@ -346,6 +348,15 @@ export interface SchoolsResponse {
export interface SchoolDetailsResponse {
school_info: School;
/**
* The published location-layer pages containing this school, widest first.
*
* Optional because the frontend and backend ship as separate images: a
* frontend deployed ahead of the API that serves this must render without
* it, not throw. Empty is also a real answer — a school whose town and
* authority both fall below the publish threshold has nowhere to link.
*/
places?: SchoolPlace[];
yearly_data: SchoolResult[];
absence_data: AbsenceData | null;
// Supplementary data (null until Kestra populates)