Compare commits

..
Author SHA1 Message Date
TudorandClaude Fable 5 d677b54533 fix(pipeline): cache invalidation for IDACI DAG too; curl timeouts
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 9m37s
PR Checks / Backend Smoke (pull_request) Successful in 6s
PR Checks / Build Backend (no push) (pull_request) Successful in 18s
PR Checks / Build Frontend (no push) (pull_request) Successful in 48s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 35s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 1m40s
Addresses AI-review findings: the annual IDACI DAG also rebuilds a mart
(fact_deprivation) and needs the reload; curl gets connect/max timeouts
so an unreachable backend fails fast instead of hanging the task.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-09 21:26:45 +01:00
TudorandClaude Fable 5 c353e36072 fix(pipeline): actually invalidate the backend cache after data rebuilds
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 9m40s
PR Checks / Backend Smoke (pull_request) Successful in 6s
PR Checks / Build Backend (no push) (pull_request) Successful in 16s
PR Checks / Build Frontend (no push) (pull_request) Successful in 51s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 36s
PR Checks / AI Code Review (Claude) (pull_request) Failing after 1m20s
The daily/monthly/annual DAG docstring promised an Invalidate Cache step
that never existed — after a marts rebuild the backend kept serving its
startup-cached (possibly empty) DataFrame until a container restart.
Add a POST /api/admin/reload task at the end of each pipeline DAG,
mirroring the sitemap DAG's admin-call pattern.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-09 21:12:01 +01:00
9 changed files with 59 additions and 45 deletions
-3
View File
@@ -78,7 +78,6 @@ OFFICIAL_SIXTH_FORM: dict[int, str] = {
0: "Not applicable", 0: "Not applicable",
1: "Has a sixth form", 1: "Has a sixth form",
2: "Does not have a sixth form", 2: "Does not have a sixth form",
9: "",
} }
RELIGIOUS_CHARACTER: dict[int, str] = { RELIGIOUS_CHARACTER: dict[int, str] = {
@@ -129,14 +128,12 @@ RELIGIOUS_CHARACTER: dict[int, str] = {
47: "Reformed Baptist", 47: "Reformed Baptist",
48: "Roman Catholic/Anglican", 48: "Roman Catholic/Anglican",
49: "Sunni Deobandi", 49: "Sunni Deobandi",
99: "",
} }
ADMISSIONS_POLICY: dict[int, str] = { ADMISSIONS_POLICY: dict[int, str] = {
0: "Not applicable", 0: "Not applicable",
2: "Selective", 2: "Selective",
4: "Non-selective", 4: "Non-selective",
9: "",
} }
-12
View File
@@ -81,15 +81,3 @@ def test_seed_matches_dictionaries():
for row in csv.DictReader(fh): for row in csv.DictReader(fh):
seed[row["field"]][int(row["code"])] = row["name"] seed[row["field"]][int(row["code"])] = row["name"]
assert seed == fields assert seed == fields
def test_blank_name_sentinel_codes_map_to_empty_string():
"""GIAS carries codes whose (name) column is blank — e.g. ReligiousCharacter
99 (~4k schools) and AdmissionsPolicy 9 (~5.6k schools). The old name
pipeline served these as empty strings; the dictionaries must reproduce
that ("" is falsy, so UI tag heuristics stay silent) rather than letting
them hit the "Unknown (<code>)" path meant for genuinely new codes."""
assert RELIGIOUS_CHARACTER[99] == ""
assert ADMISSIONS_POLICY[9] == ""
assert translate(99, RELIGIOUS_CHARACTER) == ""
assert translate(9, ADMISSIONS_POLICY) == ""
@@ -89,7 +89,7 @@ These are strengths the fixes below must not regress:
- Uplift: **home 46% exit rate — decrease, moderate** and **share of sessions reaching a school page — increase, moderate** (assists the majority entry path at its first interaction). - Uplift: **home 46% exit rate — decrease, moderate** and **share of sessions reaching a school page — increase, moderate** (assists the majority entry path at its first interaction).
- **P1.7 — The mobile hero omits the value proposition entirely** *(J1-F6)* - **P1.7 — The mobile hero omits the value proposition entirely** *(J1-F6)*
- Evidence: desktop shows the "UPDATED WITH 2026/2027 ADMISSIONS RESULTS" trust badge and the "27,000+ schools… side by side, in one place" subheading; mobile renders only the poetic H1 ("Every school in England, *compared.*") and a bare search box (`j1-home-desktop-fold.png` vs `j1-home-mobile-fold.png`). - Evidence: desktop shows the "UPDATED WITH 2026/2027 ADMISSIONS RESULTS" trust badge and the "24,000+ schools… side by side, in one place" subheading; mobile renders only the poetic H1 ("Every school in England, *compared.*") and a bare search box (`j1-home-desktop-fold.png` vs `j1-home-mobile-fold.png`).
- Criterion: mobile content parity; Nielsen #1 — first-visit orientation ("what is this, why trust it") absent on the primary viewport. - Criterion: mobile content parity; Nielsen #1 — first-visit orientation ("what is this, why trust it") absent on the primary viewport.
- Argument: 63% of entries land here and 56% of traffic is mobile; a first-time visitor gets no statement of coverage, data source, or freshness above the fold. Weak value proposition at first glance is a classic bounce driver and plausibly a material slice of the 46% exit rate. - Argument: 63% of entries land here and 56% of traffic is mobile; a first-time visitor gets no statement of coverage, data source, or freshness above the fold. Weak value proposition at first glance is a classic bounce driver and plausibly a material slice of the 46% exit rate.
- Recommendation: restore a compact version of the badge + one-line value prop under the mobile H1 (one text block; the fold has room above the deadline rail). - Recommendation: restore a compact version of the badge + one-line value prop under the mobile H1 (one text block; the fold has room above the deadline rail).
+2 -2
View File
@@ -271,10 +271,10 @@ export function HomeView({ initialSchools, filters, totalSchools, howItWorks, ed
freshness, standing in for the hidden eyebrow too) on phones, freshness, standing in for the hidden eyebrow too) on phones,
where every line above the fold costs. */} where every line above the fold costs. */}
<span className={styles.heroDescriptionFull}> <span className={styles.heroDescriptionFull}>
<strong>27,000+ primary and secondary schools</strong> with Key Stage 2 SATs, GCSE results, Ofsted grades, progress scores and admissions data side by side, in one place. <strong>24,000+ primary and secondary schools</strong> with Key Stage 2 SATs, GCSE results, Ofsted grades, progress scores and admissions data side by side, in one place.
</span> </span>
<span className={styles.heroDescriptionCompact}> <span className={styles.heroDescriptionCompact}>
<strong>27,000+ English schools</strong> SATs, GCSEs, Ofsted &amp; admissions, side by side. Updated for 2026/27. <strong>24,000+ English schools</strong> SATs, GCSEs, Ofsted &amp; admissions, side by side. Updated for 2026/27.
</span> </span>
</p> </p>
</div> </div>
+49 -4
View File
@@ -38,6 +38,31 @@ default_args = {
"retry_delay": timedelta(minutes=5), "retry_delay": timedelta(minutes=5),
} }
# The backend caches the marts DataFrame at startup; after any rebuild the
# cache must be invalidated or the API serves stale (or empty) data until the
# container restarts.
INVALIDATE_CACHE_CMD = """
set -e
BACKEND_URL="${BACKEND_URL:-http://backend:80}"
ADMIN_KEY="${ADMIN_API_KEY:-changeme}"
echo "Calling $BACKEND_URL/api/admin/reload ..."
response=$(curl -s -o /tmp/reload_response.json -w "%{http_code}" \\
--connect-timeout 10 --max-time 120 \\
-X POST "$BACKEND_URL/api/admin/reload" \\
-H "X-API-Key: $ADMIN_KEY" \\
-H "Content-Type: application/json")
echo "HTTP status: $response"
cat /tmp/reload_response.json
if [ "$response" != "200" ]; then
echo "ERROR: backend cache reload failed (HTTP $response)"
exit 1
fi
"""
# ── Daily DAG (GIAS + downstream) ────────────────────────────────────── # ── Daily DAG (GIAS + downstream) ──────────────────────────────────────
@@ -91,7 +116,12 @@ print(f'Validation passed: {{count}} GIAS rows')
bash_command=f"cd {PIPELINE_DIR} && python scripts/sync_typesense.py", bash_command=f"cd {PIPELINE_DIR} && python scripts/sync_typesense.py",
) )
extract_group >> validate_raw >> dbt_build >> sync_typesense invalidate_cache = BashOperator(
task_id="invalidate_cache",
bash_command=INVALIDATE_CACHE_CMD,
)
extract_group >> validate_raw >> dbt_build >> sync_typesense >> invalidate_cache
# ── Monthly DAG (Ofsted) ─────────────────────────────────────────────── # ── Monthly DAG (Ofsted) ───────────────────────────────────────────────
@@ -121,7 +151,12 @@ with DAG(
bash_command=f"cd {PIPELINE_DIR} && python scripts/sync_typesense.py", bash_command=f"cd {PIPELINE_DIR} && python scripts/sync_typesense.py",
) )
extract_ofsted >> dbt_build_ofsted >> sync_typesense_ofsted invalidate_cache_ofsted = BashOperator(
task_id="invalidate_cache",
bash_command=INVALIDATE_CACHE_CMD,
)
extract_ofsted >> dbt_build_ofsted >> sync_typesense_ofsted >> invalidate_cache_ofsted
# ── Annual DAG (EES: KS2, KS4, Census, Admissions) ─────────────────── # ── Annual DAG (EES: KS2, KS4, Census, Admissions) ───────────────────
@@ -153,7 +188,12 @@ with DAG(
bash_command=f"cd {PIPELINE_DIR} && python scripts/sync_typesense.py", bash_command=f"cd {PIPELINE_DIR} && python scripts/sync_typesense.py",
) )
extract_ees_group >> dbt_build_ees >> sync_typesense_ees invalidate_cache_ees = BashOperator(
task_id="invalidate_cache",
bash_command=INVALIDATE_CACHE_CMD,
)
extract_ees_group >> dbt_build_ees >> sync_typesense_ees >> invalidate_cache_ees
# ── Annual DAG (IDACI Deprivation) ──────────────────────────────────── # ── Annual DAG (IDACI Deprivation) ────────────────────────────────────
@@ -178,4 +218,9 @@ with DAG(
bash_command=f"cd {PIPELINE_DIR}/transform && {DBT_BIN} build --profiles-dir . --target production --select stg_idaci+ fact_deprivation+", bash_command=f"cd {PIPELINE_DIR}/transform && {DBT_BIN} build --profiles-dir . --target production --select stg_idaci+ fact_deprivation+",
) )
extract_idaci >> dbt_build_idaci invalidate_cache_idaci = BashOperator(
task_id="invalidate_cache",
bash_command=INVALIDATE_CACHE_CMD,
)
extract_idaci >> dbt_build_idaci >> invalidate_cache_idaci
+5 -15
View File
@@ -94,23 +94,13 @@ def main() -> None:
for code_col, name_col, dict_name, field_key in FIELDS: for code_col, name_col, dict_name, field_key in FIELDS:
pairs = ( pairs = (
df[[code_col, name_col]] df[[code_col, name_col]]
.loc[lambda d: d[code_col] != ""] .loc[lambda d: (d[code_col] != "") & (d[name_col] != "")]
.drop_duplicates() .drop_duplicates()
) )
by_code: dict[int, set] = {} mapping = sorted((int(c), n) for c, n in pairs.itertuples(index=False))
for c, n in pairs.itertuples(index=False): dupes = len(mapping) - len({c for c, _ in mapping})
by_code.setdefault(int(c), set()).add(n) if dupes:
mapping = [] sys.exit(f"{code_col}: {dupes} codes map to multiple names — investigate before generating")
for code, names in sorted(by_code.items()):
named = sorted(n for n in names if n != "")
if len(named) > 1:
sys.exit(f"{code_col}: code {code} maps to multiple names {named} — investigate before generating")
# Codes that only ever appear with a blank (name) are GIAS
# "not recorded" sentinels (e.g. ReligiousCharacter 99,
# AdmissionsPolicy 9). Map them to "" so the API serves the same
# empty string the old name pipeline did — the "Unknown (<code>)"
# path is reserved for genuinely new codes.
mapping.append((code, named[0] if named else ""))
lines = [f"{dict_name}: dict[int, str] = {{"] lines = [f"{dict_name}: dict[int, str] = {{"]
for code, name in mapping: for code, name in mapping:
escaped = name.replace('"', '\\"') escaped = name.replace('"', '\\"')
-3
View File
@@ -78,7 +78,6 @@ OFFICIAL_SIXTH_FORM: dict[int, str] = {
0: "Not applicable", 0: "Not applicable",
1: "Has a sixth form", 1: "Has a sixth form",
2: "Does not have a sixth form", 2: "Does not have a sixth form",
9: "",
} }
RELIGIOUS_CHARACTER: dict[int, str] = { RELIGIOUS_CHARACTER: dict[int, str] = {
@@ -129,14 +128,12 @@ RELIGIOUS_CHARACTER: dict[int, str] = {
47: "Reformed Baptist", 47: "Reformed Baptist",
48: "Roman Catholic/Anglican", 48: "Roman Catholic/Anglican",
49: "Sunni Deobandi", 49: "Sunni Deobandi",
99: "",
} }
ADMISSIONS_POLICY: dict[int, str] = { ADMISSIONS_POLICY: dict[int, str] = {
0: "Not applicable", 0: "Not applicable",
2: "Selective", 2: "Selective",
4: "Non-selective", 4: "Non-selective",
9: "",
} }
@@ -42,12 +42,12 @@ models:
tests: tests:
- accepted_values: - accepted_values:
severity: warn severity: warn
values: [0, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 24, 25, 26, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 99] values: [0, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 24, 25, 26, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49]
- name: admissions_policy_code - name: admissions_policy_code
tests: tests:
- accepted_values: - accepted_values:
severity: warn severity: warn
values: [0, 2, 4, 9] values: [0, 2, 4]
- name: dim_location - name: dim_location
description: School location dimension with PostGIS geometry description: School location dimension with PostGIS geometry
@@ -53,7 +53,6 @@ phase_of_education,7,All-through
official_sixth_form,0,Not applicable official_sixth_form,0,Not applicable
official_sixth_form,1,Has a sixth form official_sixth_form,1,Has a sixth form
official_sixth_form,2,Does not have a sixth form official_sixth_form,2,Does not have a sixth form
official_sixth_form,9,
religious_character,0,Does not apply religious_character,0,Does not apply
religious_character,2,Church of England religious_character,2,Church of England
religious_character,3,Roman Catholic religious_character,3,Roman Catholic
@@ -101,8 +100,6 @@ religious_character,46,Protestant/Evangelical
religious_character,47,Reformed Baptist religious_character,47,Reformed Baptist
religious_character,48,Roman Catholic/Anglican religious_character,48,Roman Catholic/Anglican
religious_character,49,Sunni Deobandi religious_character,49,Sunni Deobandi
religious_character,99,
admissions_policy,0,Not applicable admissions_policy,0,Not applicable
admissions_policy,2,Selective admissions_policy,2,Selective
admissions_policy,4,Non-selective admissions_policy,4,Non-selective
admissions_policy,9,
1 field code name
53 official_sixth_form 0 Not applicable
54 official_sixth_form 1 Has a sixth form
55 official_sixth_form 2 Does not have a sixth form
official_sixth_form 9
56 religious_character 0 Does not apply
57 religious_character 2 Church of England
58 religious_character 3 Roman Catholic
100 religious_character 47 Reformed Baptist
101 religious_character 48 Roman Catholic/Anglican
102 religious_character 49 Sunni Deobandi
religious_character 99
103 admissions_policy 0 Not applicable
104 admissions_policy 2 Selective
105 admissions_policy 4 Non-selective
admissions_policy 9