fix(places): address review, and merge places GIAS spells more than one way
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m2s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 17s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 2m34s

Two findings from review on #121, plus a third the review prompted.

The cap at three authorities silently dropped the fourth in exactly the case
where the information matters most — a genuinely fragmented place — and
contradicted the stated goal of naming every authority a place sits in. It is
gone. The share rule was always the real limit and already bounds the list at
ten. Measured against the live corpus, one town would have been truncated
today: LONDON, split evenly between Hackney, Lambeth, Westminster and
Lewisham.

parent_authority used mode() while authorities used value_counts(), and on an
exact tie pandas does not guarantee the two pick the same name, so the 301
could have pointed somewhere other than the authority named first on the page.
The parent is now derived from authorities[0]: one computation, one answer.
It also inherits the sentinel filter, so a place can no longer redirect to
/schools/authority/does-not-apply.

Chasing the truncation case surfaced a worse bug. Places were grouped by raw
town value, but the registry is keyed by slug, and GIAS spells the same place
several ways. Five town slugs come from more than one spelling: "London"
(1,819 schools) and "LONDON" (12) both slugify to `london`, so the later group
simply overwrote the earlier one — /schools/london could have shown twelve
schools, silently, depending on row order. Weston-super-Mare was split 14/19
across two spellings and Newcastle-under-Lyme across three. Grouping is now by
slug, and the display name is the most common spelling.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
This commit is contained in:
TudorandClaude Opus 5 committed 2026-08-21 22:42:18 +01:00
1 parent 1cb5314c53
commit bb2f7a5841
2 files changed
+102 -23

No files matched your search

+53
View File
@@ -290,3 +290,56 @@ def test_a_place_always_names_at_least_one_authority():
place = reg.get("town:fragmented")
assert place is not None
assert len(place.authorities) == 1
def test_every_qualifying_authority_is_named_with_no_cap():
"""An earlier cut stopped at three, dropping the fourth silently.
That truncation bit exactly where the information matters most — a
genuinely fragmented place — and nothing recorded it.
"""
rows = []
for i, la in enumerate(["Hackney", "Lambeth", "Westminster", "Lewisham"]):
rows += _town(3, "Fourway", la, start=300000 + i * 100)
reg = build_place_registry(_df(rows))
assert len(reg["town:fourway"].authorities) == 4
def test_the_redirect_target_is_the_authority_named_first():
"""They were computed separately — mode() against value_counts() — and on
an exact tie pandas does not guarantee the two agree."""
rows = (_town(26, "London", "Merton", start=300000)
+ _town(7, "London", "Wandsworth", start=400000))
for r in rows:
r["postcode"] = "SW19 1AA"
place = build_place_registry(_df(rows))["outcode:sw19"]
assert place.parent_authority == place.authorities[0][0]
def test_a_place_never_redirects_to_a_sentinel_authority():
# Deriving the parent from `authorities` inherits its sentinel filter.
rows = (_town(6, "Someplace", "Does not apply", start=300000)
+ _town(5, "Someplace", "Essex", start=400000))
reg = build_place_registry(_df(rows))
assert reg["town:someplace"].parent_authority == "Essex"
def test_spellings_of_one_place_are_merged_not_overwritten():
"""GIAS spells the same place several ways, and they share a URL.
"London" (1,819 schools) and "LONDON" (12) both slugify to `london`.
Grouping by raw value let the later group overwrite the earlier one, so
the page could have shown twelve schools instead of 1,819 — silently, and
depending on row order.
"""
rows = (_town(6, "Weston-super-Mare", "North Somerset", start=300000)
+ _town(5, "Weston-Super-Mare", "North Somerset", start=400000))
reg = build_place_registry(_df(rows))
assert len(reg["town:weston-super-mare"].urns) == 11
def test_the_merged_place_takes_its_most_common_spelling():
rows = (_town(9, "Newcastle-under-Lyme", "Staffordshire", start=300000)
+ _town(5, "NEWCASTLE-UNDER-LYME", "Staffordshire", start=400000))
reg = build_place_registry(_df(rows))
assert reg["town:newcastle-under-lyme"].name == "Newcastle-under-Lyme"