feat(admissions): show the last distance offered where councils publish it
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 7s
PR Checks / Build Backend (no push) (pull_request) Successful in 31s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 1m12s
PR Checks / AI Code Review (Claude) (pull_request) Failing after 2m7s
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 7s
PR Checks / Build Backend (no push) (pull_request) Successful in 31s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 1m12s
PR Checks / AI Code Review (Claude) (pull_request) Failing after 2m7s
Adds the cut-off distance a parent actually asks about — "how close do we
need to live?" — end to end: a Singer tap, dbt staging and mart models, an
Airflow DAG, and a tile on both detail templates. 3,597 schools across 57
local authorities carry a figure; the rest are unchanged.
There is no national source for this. Each LA publishes its own cut-offs in
its own format, and the collected CSV is transcribed from PDFs, spreadsheets
and web pages — so most of the work here is deciding what is safe to show.
Data
* tap-uk-school-distance loads the CSV verbatim into raw. Keyed on
(urn, year, school_name), because school_name carries the admission
route: (urn, year) alone collides on 118 keys and a reload would have
silently dropped every band but one.
* stg_school_distance applies a 25 m – 25 km plausibility band. The source
contains 0.0-mile rows (published where a school filled on a higher
criterion), 1-metre cut-offs, and one reading 533 miles — ~4% of rows,
all of which would put a visibly wrong number on a live page.
* fact_admission_distance collapses routes to one row per school per year
using the furthest, and keeps route_count so the page can say the figure
is the widest of several bands rather than the one for a given child.
Serving
* Kept out of fact_admissions: that mart is EES-derived and near-complete
for England, this one covers 57 LAs, and the two refresh independently.
* Latest year only. Coverage is ragged — a school may have 2021 and 2026
and nothing between — so a history array would invite a trend line drawn
through gaps that are absences of publication, not of a cut-off.
* The Admissions section now renders on either source. 3% of the schools
that render have a cut-off and no EES admissions row, and gating on
admissions alone would have hidden the figure on those pages.
Interface
* The year travels with the figure everywhere it appears; a cut-off
detached from its admissions round is not a fact about anything.
* "Not a fixed catchment — it moves every year" sits under every instance,
because that is the inference a parent will otherwise draw.
* Replaces a hardcoded "Historical distance cut-off data is not available
for this school" that appeared on every secondary page, including the
ones whose council does publish it. The absence is now stated only when
it is real, and names the authority that would hold it.
The tint costs the muted tokens their AA margin: measured on the composited
backdrop (not the computed one, which reports the untinted card), --text-muted
falls to 4.09:1 in dark theme. The tile uses --text-secondary instead — 6.50:1
dark, 6.60:1 light.
The DAG is manual, like the other annual ones: councils publish on allocation
day, each on its own timetable, so there is no date worth scheduling against.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WDvkyqqHABm4bmth2kjAxE
This commit is contained in:
1 parent
5156a85bd1
commit
88c653215d
31 files changed
+1753
-42
No files matched your search
@@ -0,0 +1,16 @@
|
||||
[build-system]
|
||||
requires = ["setuptools>=68", "wheel"]
|
||||
build-backend = "setuptools.build_meta"
|
||||
|
||||
[project]
|
||||
name = "tap-uk-school-distance"
|
||||
version = "0.1.0"
|
||||
description = "Singer tap for last-distance-offered admission cut-offs collected from LA admissions publications"
|
||||
requires-python = ">=3.10"
|
||||
dependencies = [
|
||||
"singer-sdk~=0.53",
|
||||
"requests>=2.31",
|
||||
]
|
||||
|
||||
[project.scripts]
|
||||
tap-uk-school-distance = "tap_uk_school_distance.tap:TapUKSchoolDistance.cli"
|
||||
@@ -0,0 +1 @@
|
||||
"""tap-uk-school-distance: Singer tap for LA "last distance offered" admission cut-offs."""
|
||||
@@ -0,0 +1,135 @@
|
||||
"""Singer tap for "last distance offered" school admission cut-offs.
|
||||
|
||||
Unlike every other tap in this project, the upstream here is not a single
|
||||
government endpoint. There is no national dataset of admission cut-off
|
||||
distances: each local authority publishes its own, in its own format
|
||||
(PDF booklets, XLSX allocation tables, HTML pages), and only some publish
|
||||
at all. The CSV this tap reads is the output of that collection work —
|
||||
one row per school per year per admission route, already normalised to
|
||||
metres in `distance_m`.
|
||||
|
||||
The tap deliberately does no cleaning. Implausible values (0.0 miles,
|
||||
1-metre cut-offs, a 533-mile outlier) are loaded verbatim into `raw` and
|
||||
filtered in `stg_school_distance`, so the raw table stays a faithful
|
||||
record of what the councils published and the plausibility rules live in
|
||||
one reviewable place.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import csv
|
||||
import io
|
||||
|
||||
import requests
|
||||
from singer_sdk import Stream, Tap
|
||||
from singer_sdk import typing as th
|
||||
|
||||
REQUEST_TIMEOUT = 300
|
||||
|
||||
|
||||
class SchoolDistanceStream(Stream):
|
||||
"""Stream: last distance offered, one row per school × year × admission route."""
|
||||
|
||||
name = "school_distance_offered"
|
||||
# school_name is part of the key, not decoration: where a school admits
|
||||
# through several routes the route is encoded in the name ("… Band 1",
|
||||
# "… Band 2"), and (urn, year) alone is not unique. Verified against the
|
||||
# source: (urn, year) collides on 118 keys, (urn, year, school_name) on
|
||||
# none. Without school_name in the key a reload would silently drop the
|
||||
# other bands.
|
||||
primary_keys = ["urn", "year", "school_name"]
|
||||
replication_key = None
|
||||
|
||||
schema = th.PropertiesList(
|
||||
th.Property("urn", th.IntegerType, required=True),
|
||||
th.Property("year", th.IntegerType, required=True),
|
||||
th.Property("school_name", th.StringType, required=True),
|
||||
th.Property("la_code", th.IntegerType),
|
||||
th.Property("la_name", th.StringType),
|
||||
# The figure as the council published it, before unit conversion —
|
||||
# kept so a surprising metre value can be traced back to "0.31 miles"
|
||||
# without reopening the source PDF.
|
||||
th.Property("distance_value_raw", th.NumberType),
|
||||
th.Property("distance_unit_raw", th.StringType),
|
||||
th.Property("distance_m", th.NumberType),
|
||||
# Path of the council publication the row was read from. This is the
|
||||
# provenance trail for a figure we cannot re-derive from an API.
|
||||
th.Property("source_file", th.StringType),
|
||||
).to_dict()
|
||||
|
||||
@staticmethod
|
||||
def _num(value: str) -> float | None:
|
||||
try:
|
||||
return float(value)
|
||||
except (TypeError, ValueError):
|
||||
return None
|
||||
|
||||
@staticmethod
|
||||
def _int(value: str) -> int | None:
|
||||
try:
|
||||
return int(float(value))
|
||||
except (TypeError, ValueError):
|
||||
return None
|
||||
|
||||
def get_records(self, context):
|
||||
url = self.config["download_url"]
|
||||
self.logger.info("Downloading school distance data from %s", url)
|
||||
|
||||
resp = requests.get(url, timeout=REQUEST_TIMEOUT)
|
||||
resp.raise_for_status()
|
||||
|
||||
reader = csv.DictReader(io.StringIO(resp.text))
|
||||
|
||||
yielded = 0
|
||||
skipped = 0
|
||||
for row in reader:
|
||||
urn = self._int(row.get("urn"))
|
||||
year = self._int(row.get("year"))
|
||||
name = (row.get("school_name") or "").strip()
|
||||
|
||||
# Every part of the primary key must be present, or the load is
|
||||
# not idempotent. One row in the source has a blank year.
|
||||
if urn is None or year is None or not name:
|
||||
skipped += 1
|
||||
continue
|
||||
|
||||
yield {
|
||||
"urn": urn,
|
||||
"year": year,
|
||||
"school_name": name,
|
||||
"la_code": self._int(row.get("la_code")),
|
||||
"la_name": (row.get("la_name") or "").strip() or None,
|
||||
"distance_value_raw": self._num(row.get("distance_value_raw")),
|
||||
"distance_unit_raw": (row.get("distance_unit_raw") or "").strip() or None,
|
||||
"distance_m": self._num(row.get("distance_m")),
|
||||
"source_file": (row.get("source_file") or "").strip() or None,
|
||||
}
|
||||
yielded += 1
|
||||
|
||||
self.logger.info(
|
||||
"Yielded %d distance records (%d skipped: missing urn, year or school name)",
|
||||
yielded,
|
||||
skipped,
|
||||
)
|
||||
|
||||
|
||||
class TapUKSchoolDistance(Tap):
|
||||
"""Singer tap for LA-published admission cut-off distances."""
|
||||
|
||||
name = "tap-uk-school-distance"
|
||||
|
||||
config_jsonschema = th.PropertiesList(
|
||||
th.Property(
|
||||
"download_url",
|
||||
th.StringType,
|
||||
required=True,
|
||||
description="URL of the collected last-distance-offered CSV",
|
||||
),
|
||||
).to_dict()
|
||||
|
||||
def discover_streams(self):
|
||||
return [SchoolDistanceStream(self)]
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
TapUKSchoolDistance.cli()
|
||||
Reference in new issue
Block a user