diff --git a/docs/superpowers/plans/2026-07-13-compare-api-enrichment.md b/docs/superpowers/plans/2026-07-13-compare-api-enrichment.md new file mode 100644 index 0000000..4a55d83 --- /dev/null +++ b/docs/superpowers/plans/2026-07-13-compare-api-enrichment.md @@ -0,0 +1,399 @@ +# Compare API Enrichment (Backend PR) Implementation Plan + +> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking. + +**Goal:** Expose the PR #32 data through the API so the redesigned compare screen can be built: enrich `/api/compare` with supplementary blocks + national averages + computed benchmarks, translate Ofsted report-card codes to labels, and surface the new mart columns (spec §6, §8 of `docs/superpowers/specs/2026-07-11-compare-screen-redesign-design.md`). + +**Architecture:** All changes are additive API fields — existing consumers keep working. One small dbt change rides along: `fact_performance` (the combined KS2+KS4 mart the backend's `_MAIN_QUERY` reads) enumerates columns explicitly and was not extended in PR #32, so the new KS2 CI and KS4 banding columns must be threaded through it here. Everything else is backend Python: `models.py` mappings, `data_loader` query/supplementary additions, an Ofsted label dictionary (gias_codes pattern), and `/api/compare` composition. + +**Tech Stack:** FastAPI, SQLAlchemy, pandas; dbt (one model); pytest via `python -m pytest backend/tests -q` (CI installs `requirements.txt pytest "httpx<0.28"`; locally use `uv run --with-requirements requirements.txt --with pytest --with "httpx==0.27.0" python -m pytest backend/tests -q`). + +## Global Constraints + +- **Never push to `main`.** Branch: `feat/compare-api-enrichment`. +- **Additive only** to API responses; never rename/remove existing fields (frontend + e2e depend on them). +- **Report-card scale labels are the live-sampled vocabulary** (evidence in `pipeline/scripts/diagnose_compare_gaps.py`): `1=Exceptional, 2=Strong standard, 3=Expected standard, 4=Needs attention, 5=Urgent improvement`. Never "Attention needed". Safeguarding is boolean met/not-met, never counted as a graded area. +- **Ofsted links** are always the provider page `https://reports.ofsted.gov.uk/provider/21/{urn}` (spec §5) labelled as the school's Ofsted page. +- **Benchmark provenance** (spec §8.6): computed values are "state-school average (computed from our dataset)" — the API must expose them under a `benchmarks` key, clearly separate from official `national_averages`. +- TDD: each behaviour lands with a failing test first, in `backend/tests/` following the `test_school_details.py` pattern (pandas fixture + monkeypatched `load_school_data` + `TestClient`). +- Deploy note for the PR body: the new API fields return NULL/empty until prod's DAGs have run post-promotion. + +--- + +### Task 0: Branch + +- [ ] `git checkout main && git pull && git checkout -b feat/compare-api-enrichment` (commit this plan file on the branch). + +--- + +### Task 1: Thread PR #32 columns through `fact_performance` + +**Files:** +- Modify: `pipeline/transform/models/marts/fact_performance.sql` +- Modify: `pipeline/transform/models/marts/_marts_schema.yml` (fact_performance block, if it has one — add the columns wherever the model's other columns are listed; if the model has no column list there, skip the yml) + +**Interfaces:** +- Produces (for `_MAIN_QUERY` in Task 4): `ks2.*` CI columns and `ks4.progress_8_banding`, `ks4.attainment_8_disadvantage_gap`, `ks4.progress_8_disadvantage_gap` on `marts.fact_performance`. + +- [ ] **Step 1:** In `fact_performance.sql`, after `ks2.reading_progress,` add `ks2.reading_progress_lower_ci,` and `ks2.reading_progress_upper_ci,`; after `ks2.writing_progress,` add `ks2.writing_progress_lower_ci,`, `ks2.writing_progress_upper_ci,`, `ks2.writing_working_towards_pct,`; after `ks2.maths_progress,` add `ks2.maths_progress_lower_ci,`, `ks2.maths_progress_upper_ci,`. In the KS4 section, after the `ks4.progress_8_upper_ci`-equivalent line (locate the Progress 8 block) add: + +```sql + ks4.progress_8_banding, + ks4.attainment_8_disadvantage_gap, + ks4.progress_8_disadvantage_gap, +``` + +- [ ] **Step 2:** Parse gate: `cd pipeline/transform && uv run --with dbt-postgres python -m dbt.cli.main parse --profiles-dir .` → exit 0. + +- [ ] **Step 3:** Commit: `feat(pipeline): thread compare-foundation columns through fact_performance` + +--- + +### Task 2: ORM mappings for the new mart columns + +**Files:** +- Modify: `backend/models.py` (`KS2Performance` after `maths_progress`; `FactAdmissions` after `first_preference_offers`) +- Test: none (declarative mappings; covered by Task 4's query tests) + +**Interfaces:** +- Produces attributes used by Task 4: `KS2Performance.reading_progress_lower_ci` … `maths_progress_upper_ci`, `writing_working_towards_pct` (Float); `FactAdmissions.total_offers`, `.second_preference_offers`, `.third_preference_offers`, `.cross_la_applications`, `.cross_la_offers` (Integer). + +- [ ] **Step 1:** Add to `KS2Performance` (next to the existing progress columns): + +```python + reading_progress_lower_ci = Column(Float) + reading_progress_upper_ci = Column(Float) + writing_progress_lower_ci = Column(Float) + writing_progress_upper_ci = Column(Float) + writing_working_towards_pct = Column(Float) + maths_progress_lower_ci = Column(Float) + maths_progress_upper_ci = Column(Float) +``` + +Add to `FactAdmissions` (after `first_preference_offers`): + +```python + total_offers = Column(Integer) + second_preference_offers = Column(Integer) + third_preference_offers = Column(Integer) + cross_la_applications = Column(Integer) + cross_la_offers = Column(Integer) +``` + +(`FactOfstedInspection` already maps all `rc_*` columns with the right types — verify, don't change.) + +- [ ] **Step 2:** Commit: `feat(api): map compare-foundation mart columns` + +--- + +### Task 3: Ofsted label dictionary + provider URL (TDD) + +**Files:** +- Create: `backend/ofsted_codes.py` +- Test: `backend/tests/test_ofsted_codes.py` + +**Interfaces:** +- Produces for Task 4: `REPORT_CARD_GRADE_NAMES: dict[int, str]`, `report_card_labels(ofsted: dict) -> dict` (returns `{area_key: {"code": int, "label": str}}` for the non-null `rc_*` grade fields, excluding safeguarding), `ofsted_page_url(urn: int) -> str`. + +- [ ] **Step 1: Failing tests** + +```python +"""Report-card code translation uses the live-sampled Ofsted vocabulary +(pipeline/scripts/diagnose_compare_gaps.py TASK 7 VALUE SAMPLE): +Exceptional / Strong standard / Expected standard / Needs attention / +Urgent improvement — never the consultation draft's 'Attention needed'.""" +from backend.ofsted_codes import ( + REPORT_CARD_GRADE_NAMES, report_card_labels, ofsted_page_url, +) + + +def test_scale_is_sampled_vocabulary(): + assert REPORT_CARD_GRADE_NAMES == { + 1: "Exceptional", + 2: "Strong standard", + 3: "Expected standard", + 4: "Needs attention", + 5: "Urgent improvement", + } + + +def test_labels_only_for_populated_areas_and_never_safeguarding(): + ofsted = { + "rc_achievement": 2, + "rc_inclusion": 3, + "rc_attendance_behaviour": 4, + "rc_early_years": None, + "rc_safeguarding_met": True, + "overall_effectiveness": None, + } + labels = report_card_labels(ofsted) + assert labels == { + "rc_achievement": {"code": 2, "label": "Strong standard"}, + "rc_inclusion": {"code": 3, "label": "Expected standard"}, + "rc_attendance_behaviour": {"code": 4, "label": "Needs attention"}, + } + + +def test_unknown_code_is_skipped_not_crashed(): + assert report_card_labels({"rc_achievement": 9}) == {} + + +def test_provider_url(): + assert ofsted_page_url(138690) == "https://reports.ofsted.gov.uk/provider/21/138690" +``` + +- [ ] **Step 2:** Run `uv run --with-requirements requirements.txt --with pytest --with "httpx==0.27.0" python -m pytest backend/tests/test_ofsted_codes.py -q` → FAIL (module missing). + +- [ ] **Step 3: Implement `backend/ofsted_codes.py`** + +```python +"""Ofsted renewed-framework (Nov 2025) report-card code translation. + +Scale labels are the live-sampled vocabulary from the Ofsted MI file +(see pipeline/scripts/diagnose_compare_gaps.py, TASK 7 VALUE SAMPLE) — +verified against real data, not the consultation draft. +""" + +REPORT_CARD_GRADE_NAMES = { + 1: "Exceptional", + 2: "Strong standard", + 3: "Expected standard", + 4: "Needs attention", + 5: "Urgent improvement", +} + +# Graded evaluation areas only — safeguarding is a separate boolean +# judgement and must never appear in grade counts or label maps. +_RC_AREA_KEYS = ( + "rc_inclusion", + "rc_curriculum_teaching", + "rc_achievement", + "rc_attendance_behaviour", + "rc_personal_development", + "rc_leadership_governance", + "rc_early_years", + "rc_sixth_form", +) + + +def report_card_labels(ofsted: dict) -> dict: + """{area_key: {code, label}} for populated, known-valued rc_* areas.""" + out = {} + for key in _RC_AREA_KEYS: + code = ofsted.get(key) + label = REPORT_CARD_GRADE_NAMES.get(code) + if code is not None and label is not None: + out[key] = {"code": code, "label": label} + return out + + +def ofsted_page_url(urn: int) -> str: + """The school's page on ofsted.gov.uk (all its reports live there — + we never deep-link an individual report; spec §5).""" + return f"https://reports.ofsted.gov.uk/provider/21/{urn}" +``` + +- [ ] **Step 4:** Re-run the test file → 4 passed. Run the full suite (same command, `backend/tests -q`) → all pass. + +- [ ] **Step 5:** Commit: `feat(api): Ofsted report-card labels and provider-page URL` + +--- + +### Task 4: data_loader — query columns + richer supplementary blocks (TDD) + +**Files:** +- Modify: `backend/data_loader.py` (`_MAIN_QUERY` ~line 153; `get_supplementary_data` ~line 460) +- Test: `backend/tests/test_supplementary_enrichment.py` + +**Interfaces:** +- `_MAIN_QUERY` additionally selects (KS2 block, after `p.maths_progress`): `p.reading_progress_lower_ci, p.reading_progress_upper_ci, p.writing_progress_lower_ci, p.writing_progress_upper_ci, p.writing_working_towards_pct, p.maths_progress_lower_ci, p.maths_progress_upper_ci`; (KS4 block, after the Progress 8 CI columns): `p.progress_8_banding, p.attainment_8_disadvantage_gap, p.progress_8_disadvantage_gap`. Note `_MAIN_QUERY_NO_SIXTH_FORM`/`_MAIN_QUERY_LEGACY_NAMES` are string-derived from `_MAIN_QUERY` (lines 259-270) and inherit automatically — verify the assertions there still hold. +- `get_supplementary_data(db, urn)["admissions"]` rows additionally carry: `total_offers`, `second_preference_offers`, `third_preference_offers`, `cross_la_applications`, `cross_la_offers` (add to `_admissions_row`). +- `get_supplementary_data(db, urn)["ofsted"]` additionally carries: `report_card` (the `report_card_labels(...)` dict, `{}` when no rc data), `ofsted_page_url`, and `grade_source`: `"graded"` when `overall_effectiveness` came from the graded column, `"ungraded_carried_forward"` when the fallback `ungraded_grade` supplied it, `None` when neither. + +- [ ] **Step 1: Failing tests** — construct a fake Ofsted row object (simple `types.SimpleNamespace` with the model's attributes) and call the block-building logic via `get_supplementary_data` with a stubbed session (follow how existing tests stub the db; if none do, factor the ofsted-dict construction into a pure helper `_ofsted_block(o, urn)` and test that directly — preferred): + +```python +import types +from backend.data_loader import _ofsted_block + + +def _row(**kw): + base = dict( + framework="RC", inspection_date=None, inspection_type=None, + overall_effectiveness=None, quality_of_education=None, + behaviour_attitudes=None, personal_development=None, + leadership_management=None, early_years_provision=None, + sixth_form_provision=None, ungraded_outcome=None, ungraded_grade=None, + rc_safeguarding_met=None, rc_inclusion=None, rc_curriculum_teaching=None, + rc_achievement=None, rc_attendance_behaviour=None, + rc_personal_development=None, rc_leadership_governance=None, + rc_early_years=None, rc_sixth_form=None, report_url=None, + ) + base.update(kw) + return types.SimpleNamespace(**base) + + +def test_report_card_block_and_provider_url(): + o = _row(rc_achievement=2, rc_inclusion=3, rc_safeguarding_met=True) + block = _ofsted_block(o, urn=100140) + assert block["report_card"]["rc_achievement"]["label"] == "Strong standard" + assert "rc_safeguarding_met" not in block["report_card"] + assert block["rc_safeguarding_met"] is True + assert block["ofsted_page_url"] == "https://reports.ofsted.gov.uk/provider/21/100140" + + +def test_grade_source_graded_vs_carried_forward(): + assert _ofsted_block(_row(overall_effectiveness=1), urn=1)["grade_source"] == "graded" + carried = _ofsted_block(_row(ungraded_grade=2), urn=1) + assert carried["grade_source"] == "ungraded_carried_forward" + assert carried["overall_effectiveness"] == 2 + assert _ofsted_block(_row(), urn=1)["grade_source"] is None + + +def test_admissions_row_new_fields(): + from backend.data_loader import _admissions_row_dict + a = types.SimpleNamespace( + year=202627, school_phase="Primary", places_offered=80, + total_applications=185, first_preference_applications=74, + first_preference_offers=74, first_preference_offer_pct=100.0, + oversubscription_ratio=0.925, oversubscribed=False, + total_offers=80, second_preference_offers=4, third_preference_offers=2, + cross_la_applications=12, cross_la_offers=3, + ) + d = _admissions_row_dict(a) + for k in ("total_offers", "second_preference_offers", "third_preference_offers", + "cross_la_applications", "cross_la_offers"): + assert d[k] == getattr(a, k) +``` + +- [ ] **Step 2:** Run → FAIL (helpers don't exist). + +- [ ] **Step 3: Implement.** Refactor the existing inline ofsted-dict construction in `get_supplementary_data` into a module-level `_ofsted_block(o, urn)` that produces the existing keys **unchanged** plus the three new ones (`report_card` via `report_card_labels(...)` from Task 3, `ofsted_page_url` via `ofsted_page_url(urn)`, `grade_source` per the interface rule — derived from which source supplied `overall_effectiveness`). Rename/extract the local `_admissions_row` into module-level `_admissions_row_dict(a)` and append the five new fields. Add the ten new columns to `_MAIN_QUERY` exactly as the interface lists them. `get_supplementary_data` calls both helpers; its external shape gains only additive keys. + +- [ ] **Step 4:** Full suite → all pass (existing `test_school_details.py` etc. must not break; if a fixture enumerates yearly-data columns, extend it with the new NaN columns as needed). + +- [ ] **Step 5:** Commit: `feat(api): expose progress CIs, KS4 banding/gaps, admissions detail, report-card labels` + +--- + +### Task 5: Computed benchmarks helper (TDD) + +**Files:** +- Modify: `backend/data_loader.py` (new function) +- Test: `backend/tests/test_benchmarks.py` + +**Interfaces:** +- Produces for Task 6: `compute_benchmarks(df) -> dict` — pure function over the main dataframe (latest year, state schools), shape: + +```python +{ + "source": "state-school average (computed from our dataset)", + "year": 202425, + "primary": { + "disadvantaged_rwm_expected_pct": 46.1, # weighted by eligible_pupils + "eal_pct": 22.3, # median + "sen_support_pct": 14.0, # median + "disadvantaged_pct": 24.8, # median (FSM6 proxy) + "median_pupils": 281, # median school size + }, + "secondary": { "median_pupils": 1024, "eal_pct": ..., "sen_support_pct": ..., "disadvantaged_pct": ... }, +} +``` + +- [ ] **Step 1: Failing tests** — build a small synthetic df (6 primary rows with known eligible_pupils/rwm_expected_disadvantaged_pct so the weighted average is hand-checkable; a couple of secondary rows flagged by non-null `attainment_8_score`), assert: weighted disadvantaged average matches hand computation (not the unweighted mean), medians ignore NaN, secondary block lacks the disadvantaged-RWM key, latest-year filtering (rows from an older year must not affect results), and empty df → `{}`. + +- [ ] **Step 2:** Run → FAIL. + +- [ ] **Step 3: Implement** in `data_loader.py`: + +```python +def compute_benchmarks(df: pd.DataFrame) -> dict: + """State-school benchmarks computed from our dataset (spec §5/§8.6). + These are NOT official DfE figures — consumers must label them + 'state-school average (computed from our dataset)'.""" + if df.empty or "year" not in df.columns: + return {} + latest_year = df["year"].max() + d = df[df["year"] == latest_year] + if d.empty: + return {} + is_secondary = d["attainment_8_score"].notna() if "attainment_8_score" in d.columns else pd.Series(False, index=d.index) + prim, sec = d[~is_secondary], d[is_secondary] + + def _median(sub, col): + if col not in sub.columns: + return None + v = sub[col].median() + return round(float(v), 1) if pd.notna(v) else None + + def _weighted_disadvantaged(sub): + if not {"rwm_expected_disadvantaged_pct", "eligible_pupils"} <= set(sub.columns): + return None + s = sub.dropna(subset=["rwm_expected_disadvantaged_pct", "eligible_pupils"]) + if s.empty or s["eligible_pupils"].sum() == 0: + return None + w = (s["rwm_expected_disadvantaged_pct"] * s["eligible_pupils"]).sum() / s["eligible_pupils"].sum() + return round(float(w), 1) + + def _block(sub, with_disadvantaged): + block = { + "eal_pct": _median(sub, "eal_pct"), + "sen_support_pct": _median(sub, "sen_support_pct"), + "disadvantaged_pct": _median(sub, "disadvantaged_pct"), + "median_pupils": int(sub["total_pupils"].median()) if "total_pupils" in sub.columns and pd.notna(sub["total_pupils"].median()) else None, + } + if with_disadvantaged: + block["disadvantaged_rwm_expected_pct"] = _weighted_disadvantaged(sub) + return block + + return { + "source": "state-school average (computed from our dataset)", + "year": int(latest_year), + "primary": _block(prim, with_disadvantaged=True), + "secondary": _block(sec, with_disadvantaged=False), + } +``` + +(Adapt column presence to the real df — `sen_support_pct` reaches the df via `_MAIN_QUERY`; confirm and add it there if the KS2 block doesn't already select it, mirroring Task 4's additions.) + +- [ ] **Step 4:** Full suite → pass. **Step 5:** Commit: `feat(api): computed state-school benchmarks` + +--- + +### Task 6: Enrich `/api/compare` + expose GPS/science national averages (TDD) + +**Files:** +- Modify: `backend/app.py` (`compare_schools` ~line 636; `get_national_averages` ~line 730) +- Test: `backend/tests/test_compare_enrichment.py` + +**Interfaces (response additions, all additive):** +- `/api/compare` top level gains: `"national_averages"` (same payload the `/api/national-averages` endpoint returns — extract the endpoint body into a helper `_national_averages_payload(df)` and reuse; do not duplicate the logic) and `"benchmarks"` (Task 5's `compute_benchmarks(df)`). +- Each `comparison[urn]` gains: `"ofsted"`, `"census"`, `"admissions"`, `"admissions_history"`, `"deprivation"` from `get_supplementary_data` (one `SessionLocal()` for the whole request, closed in `finally`; on exception the five keys are `None`/`[]` — mirror the detail endpoint's defensive pattern at app.py:583-590). +- `get_national_averages`' KS2 metric list gains `"gps_expected_pct", "gps_high_pct", "science_expected_pct"` so the England ticks for GPS/science flow once the data exists. + +- [ ] **Step 1: Failing tests** — monkeypatch `load_school_data` with a two-school primary df (reuse/extend the fixture style of `test_school_details.py`) and monkeypatch `get_supplementary_data` to a canned dict; assert on `TestClient(app).get("/api/compare?urns=...")`: + - response keeps the existing shape (`comparison[urn]["school_info"]["rwm_expected_pct"]` etc.), + - each school gains the five supplementary keys (canned values round-tripped), + - top-level `national_averages` and `benchmarks` present; `benchmarks["source"]` is the exact provenance string, + - a supplementary-layer exception (monkeypatched to raise) degrades to `ofsted: None` etc. with HTTP 200, + - `/api/national-averages` includes `gps_expected_pct` in the primary block when the df/national table provides it (monkeypatch the national-averages source the endpoint reads). + +- [ ] **Step 2:** Run → FAIL. **Step 3:** Implement per the interfaces. **Step 4:** Full suite → pass. + +- [ ] **Step 5:** Commit: `feat(api): compare endpoint carries supplementary blocks, national averages and benchmarks` + +--- + +### Task 7: PR + verification + +- [ ] **Step 1:** Full suite one more time + `uv run --with pyyaml python3 -c "import yaml; yaml.safe_load(open('.gitea/workflows/deploy.yml'))"` sanity is NOT needed (no workflow changes) — instead run the dbt parse gate again (Task 1 file). +- [ ] **Step 2:** Push, open PR via the Gitea API (credential-helper basic auth). PR body: the new response shapes (one JSON sketch), the reused-not-duplicated national-averages helper, the provenance rule for benchmarks, deploy note (fields NULL until prod DAGs run post-promotion), and that no e2e change is needed (no user-facing behaviour changes — the compare UI still reads the old fields; the frontend PR carries the journey updates). +- [ ] **Step 3:** After merge + staging deploy: `curl -s https://stx.schoolcompare.co.uk/api/compare?urns=138690,100140 | python3 -m json.tool | head -80` — verify the new keys and that `benchmarks.primary.disadvantaged_rwm_expected_pct` is plausible (~45-47). Verify `/api/national-averages` now carries `gps_expected_pct`/`science_expected_pct` (values or honest nulls if DfE suppresses them at national level). + +--- + +## Out of scope + +- Frontend rebuild + e2e journeys (next PR — consumes everything this PR exposes). +- `schemas.py` METRIC_DEFINITIONS additions for the trends picker (frontend PR decides which of the new columns become picker metrics). +- CI-based progress banding logic (frontend computes Above/Average/Below from the CI columns; historical years only).