Files
school_compare/docs/superpowers/plans/2026-07-13-compare-api-enrichment.md
T

22 KiB

Compare API Enrichment (Backend PR) Implementation Plan

For agentic workers: REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (- [ ]) syntax for tracking.

Goal: Expose the PR #32 data through the API so the redesigned compare screen can be built: enrich /api/compare with supplementary blocks + national averages + computed benchmarks, translate Ofsted report-card codes to labels, and surface the new mart columns (spec §6, §8 of docs/superpowers/specs/2026-07-11-compare-screen-redesign-design.md).

Architecture: All changes are additive API fields — existing consumers keep working. One small dbt change rides along: fact_performance (the combined KS2+KS4 mart the backend's _MAIN_QUERY reads) enumerates columns explicitly and was not extended in PR #32, so the new KS2 CI and KS4 banding columns must be threaded through it here. Everything else is backend Python: models.py mappings, data_loader query/supplementary additions, an Ofsted label dictionary (gias_codes pattern), and /api/compare composition.

Tech Stack: FastAPI, SQLAlchemy, pandas; dbt (one model); pytest via python -m pytest backend/tests -q (CI installs requirements.txt pytest "httpx<0.28"; locally use uv run --with-requirements requirements.txt --with pytest --with "httpx==0.27.0" python -m pytest backend/tests -q).

Global Constraints

  • Never push to main. Branch: feat/compare-api-enrichment.
  • Additive only to API responses; never rename/remove existing fields (frontend + e2e depend on them).
  • Report-card scale labels are the live-sampled vocabulary (evidence in pipeline/scripts/diagnose_compare_gaps.py): 1=Exceptional, 2=Strong standard, 3=Expected standard, 4=Needs attention, 5=Urgent improvement. Never "Attention needed". Safeguarding is boolean met/not-met, never counted as a graded area.
  • Ofsted links are always the provider page https://reports.ofsted.gov.uk/provider/21/{urn} (spec §5) labelled as the school's Ofsted page.
  • Benchmark provenance (spec §8.6): computed values are "state-school average (computed from our dataset)" — the API must expose them under a benchmarks key, clearly separate from official national_averages.
  • TDD: each behaviour lands with a failing test first, in backend/tests/ following the test_school_details.py pattern (pandas fixture + monkeypatched load_school_data + TestClient).
  • Deploy note for the PR body: the new API fields return NULL/empty until prod's DAGs have run post-promotion.

Task 0: Branch

  • git checkout main && git pull && git checkout -b feat/compare-api-enrichment (commit this plan file on the branch).

Task 1: Thread PR #32 columns through fact_performance

Files:

  • Modify: pipeline/transform/models/marts/fact_performance.sql
  • Modify: pipeline/transform/models/marts/_marts_schema.yml (fact_performance block, if it has one — add the columns wherever the model's other columns are listed; if the model has no column list there, skip the yml)

Interfaces:

  • Produces (for _MAIN_QUERY in Task 4): ks2.* CI columns and ks4.progress_8_banding, ks4.attainment_8_disadvantage_gap, ks4.progress_8_disadvantage_gap on marts.fact_performance.

  • Step 1: In fact_performance.sql, after ks2.reading_progress, add ks2.reading_progress_lower_ci, and ks2.reading_progress_upper_ci,; after ks2.writing_progress, add ks2.writing_progress_lower_ci,, ks2.writing_progress_upper_ci,, ks2.writing_working_towards_pct,; after ks2.maths_progress, add ks2.maths_progress_lower_ci,, ks2.maths_progress_upper_ci,. In the KS4 section, after the ks4.progress_8_upper_ci-equivalent line (locate the Progress 8 block) add:

    ks4.progress_8_banding,
    ks4.attainment_8_disadvantage_gap,
    ks4.progress_8_disadvantage_gap,
  • Step 2: Parse gate: cd pipeline/transform && uv run --with dbt-postgres python -m dbt.cli.main parse --profiles-dir . → exit 0.

  • Step 3: Commit: feat(pipeline): thread compare-foundation columns through fact_performance


Task 2: ORM mappings for the new mart columns

Files:

  • Modify: backend/models.py (KS2Performance after maths_progress; FactAdmissions after first_preference_offers)
  • Test: none (declarative mappings; covered by Task 4's query tests)

Interfaces:

  • Produces attributes used by Task 4: KS2Performance.reading_progress_lower_cimaths_progress_upper_ci, writing_working_towards_pct (Float); FactAdmissions.total_offers, .second_preference_offers, .third_preference_offers, .cross_la_applications, .cross_la_offers (Integer).

  • Step 1: Add to KS2Performance (next to the existing progress columns):

    reading_progress_lower_ci = Column(Float)
    reading_progress_upper_ci = Column(Float)
    writing_progress_lower_ci = Column(Float)
    writing_progress_upper_ci = Column(Float)
    writing_working_towards_pct = Column(Float)
    maths_progress_lower_ci = Column(Float)
    maths_progress_upper_ci = Column(Float)

Add to FactAdmissions (after first_preference_offers):

    total_offers = Column(Integer)
    second_preference_offers = Column(Integer)
    third_preference_offers = Column(Integer)
    cross_la_applications = Column(Integer)
    cross_la_offers = Column(Integer)

(FactOfstedInspection already maps all rc_* columns with the right types — verify, don't change.)

  • Step 2: Commit: feat(api): map compare-foundation mart columns

Task 3: Ofsted label dictionary + provider URL (TDD)

Files:

  • Create: backend/ofsted_codes.py
  • Test: backend/tests/test_ofsted_codes.py

Interfaces:

  • Produces for Task 4: REPORT_CARD_GRADE_NAMES: dict[int, str], report_card_labels(ofsted: dict) -> dict (returns {area_key: {"code": int, "label": str}} for the non-null rc_* grade fields, excluding safeguarding), ofsted_page_url(urn: int) -> str.

  • Step 1: Failing tests

"""Report-card code translation uses the live-sampled Ofsted vocabulary
(pipeline/scripts/diagnose_compare_gaps.py TASK 7 VALUE SAMPLE):
Exceptional / Strong standard / Expected standard / Needs attention /
Urgent improvement — never the consultation draft's 'Attention needed'."""
from backend.ofsted_codes import (
    REPORT_CARD_GRADE_NAMES, report_card_labels, ofsted_page_url,
)


def test_scale_is_sampled_vocabulary():
    assert REPORT_CARD_GRADE_NAMES == {
        1: "Exceptional",
        2: "Strong standard",
        3: "Expected standard",
        4: "Needs attention",
        5: "Urgent improvement",
    }


def test_labels_only_for_populated_areas_and_never_safeguarding():
    ofsted = {
        "rc_achievement": 2,
        "rc_inclusion": 3,
        "rc_attendance_behaviour": 4,
        "rc_early_years": None,
        "rc_safeguarding_met": True,
        "overall_effectiveness": None,
    }
    labels = report_card_labels(ofsted)
    assert labels == {
        "rc_achievement": {"code": 2, "label": "Strong standard"},
        "rc_inclusion": {"code": 3, "label": "Expected standard"},
        "rc_attendance_behaviour": {"code": 4, "label": "Needs attention"},
    }


def test_unknown_code_is_skipped_not_crashed():
    assert report_card_labels({"rc_achievement": 9}) == {}


def test_provider_url():
    assert ofsted_page_url(138690) == "https://reports.ofsted.gov.uk/provider/21/138690"
  • Step 2: Run uv run --with-requirements requirements.txt --with pytest --with "httpx==0.27.0" python -m pytest backend/tests/test_ofsted_codes.py -q → FAIL (module missing).

  • Step 3: Implement backend/ofsted_codes.py

"""Ofsted renewed-framework (Nov 2025) report-card code translation.

Scale labels are the live-sampled vocabulary from the Ofsted MI file
(see pipeline/scripts/diagnose_compare_gaps.py, TASK 7 VALUE SAMPLE) —
verified against real data, not the consultation draft.
"""

REPORT_CARD_GRADE_NAMES = {
    1: "Exceptional",
    2: "Strong standard",
    3: "Expected standard",
    4: "Needs attention",
    5: "Urgent improvement",
}

# Graded evaluation areas only — safeguarding is a separate boolean
# judgement and must never appear in grade counts or label maps.
_RC_AREA_KEYS = (
    "rc_inclusion",
    "rc_curriculum_teaching",
    "rc_achievement",
    "rc_attendance_behaviour",
    "rc_personal_development",
    "rc_leadership_governance",
    "rc_early_years",
    "rc_sixth_form",
)


def report_card_labels(ofsted: dict) -> dict:
    """{area_key: {code, label}} for populated, known-valued rc_* areas."""
    out = {}
    for key in _RC_AREA_KEYS:
        code = ofsted.get(key)
        label = REPORT_CARD_GRADE_NAMES.get(code)
        if code is not None and label is not None:
            out[key] = {"code": code, "label": label}
    return out


def ofsted_page_url(urn: int) -> str:
    """The school's page on ofsted.gov.uk (all its reports live there —
    we never deep-link an individual report; spec §5)."""
    return f"https://reports.ofsted.gov.uk/provider/21/{urn}"
  • Step 4: Re-run the test file → 4 passed. Run the full suite (same command, backend/tests -q) → all pass.

  • Step 5: Commit: feat(api): Ofsted report-card labels and provider-page URL


Task 4: data_loader — query columns + richer supplementary blocks (TDD)

Files:

  • Modify: backend/data_loader.py (_MAIN_QUERY ~line 153; get_supplementary_data ~line 460)
  • Test: backend/tests/test_supplementary_enrichment.py

Interfaces:

  • _MAIN_QUERY additionally selects (KS2 block, after p.maths_progress): p.reading_progress_lower_ci, p.reading_progress_upper_ci, p.writing_progress_lower_ci, p.writing_progress_upper_ci, p.writing_working_towards_pct, p.maths_progress_lower_ci, p.maths_progress_upper_ci; (KS4 block, after the Progress 8 CI columns): p.progress_8_banding, p.attainment_8_disadvantage_gap, p.progress_8_disadvantage_gap. Note _MAIN_QUERY_NO_SIXTH_FORM/_MAIN_QUERY_LEGACY_NAMES are string-derived from _MAIN_QUERY (lines 259-270) and inherit automatically — verify the assertions there still hold.

  • get_supplementary_data(db, urn)["admissions"] rows additionally carry: total_offers, second_preference_offers, third_preference_offers, cross_la_applications, cross_la_offers (add to _admissions_row).

  • get_supplementary_data(db, urn)["ofsted"] additionally carries: report_card (the report_card_labels(...) dict, {} when no rc data), ofsted_page_url, and grade_source: "graded" when overall_effectiveness came from the graded column, "ungraded_carried_forward" when the fallback ungraded_grade supplied it, None when neither.

  • Step 1: Failing tests — construct a fake Ofsted row object (simple types.SimpleNamespace with the model's attributes) and call the block-building logic via get_supplementary_data with a stubbed session (follow how existing tests stub the db; if none do, factor the ofsted-dict construction into a pure helper _ofsted_block(o, urn) and test that directly — preferred):

import types
from backend.data_loader import _ofsted_block


def _row(**kw):
    base = dict(
        framework="RC", inspection_date=None, inspection_type=None,
        overall_effectiveness=None, quality_of_education=None,
        behaviour_attitudes=None, personal_development=None,
        leadership_management=None, early_years_provision=None,
        sixth_form_provision=None, ungraded_outcome=None, ungraded_grade=None,
        rc_safeguarding_met=None, rc_inclusion=None, rc_curriculum_teaching=None,
        rc_achievement=None, rc_attendance_behaviour=None,
        rc_personal_development=None, rc_leadership_governance=None,
        rc_early_years=None, rc_sixth_form=None, report_url=None,
    )
    base.update(kw)
    return types.SimpleNamespace(**base)


def test_report_card_block_and_provider_url():
    o = _row(rc_achievement=2, rc_inclusion=3, rc_safeguarding_met=True)
    block = _ofsted_block(o, urn=100140)
    assert block["report_card"]["rc_achievement"]["label"] == "Strong standard"
    assert "rc_safeguarding_met" not in block["report_card"]
    assert block["rc_safeguarding_met"] is True
    assert block["ofsted_page_url"] == "https://reports.ofsted.gov.uk/provider/21/100140"


def test_grade_source_graded_vs_carried_forward():
    assert _ofsted_block(_row(overall_effectiveness=1), urn=1)["grade_source"] == "graded"
    carried = _ofsted_block(_row(ungraded_grade=2), urn=1)
    assert carried["grade_source"] == "ungraded_carried_forward"
    assert carried["overall_effectiveness"] == 2
    assert _ofsted_block(_row(), urn=1)["grade_source"] is None


def test_admissions_row_new_fields():
    from backend.data_loader import _admissions_row_dict
    a = types.SimpleNamespace(
        year=202627, school_phase="Primary", places_offered=80,
        total_applications=185, first_preference_applications=74,
        first_preference_offers=74, first_preference_offer_pct=100.0,
        oversubscription_ratio=0.925, oversubscribed=False,
        total_offers=80, second_preference_offers=4, third_preference_offers=2,
        cross_la_applications=12, cross_la_offers=3,
    )
    d = _admissions_row_dict(a)
    for k in ("total_offers", "second_preference_offers", "third_preference_offers",
              "cross_la_applications", "cross_la_offers"):
        assert d[k] == getattr(a, k)
  • Step 2: Run → FAIL (helpers don't exist).

  • Step 3: Implement. Refactor the existing inline ofsted-dict construction in get_supplementary_data into a module-level _ofsted_block(o, urn) that produces the existing keys unchanged plus the three new ones (report_card via report_card_labels(...) from Task 3, ofsted_page_url via ofsted_page_url(urn), grade_source per the interface rule — derived from which source supplied overall_effectiveness). Rename/extract the local _admissions_row into module-level _admissions_row_dict(a) and append the five new fields. Add the ten new columns to _MAIN_QUERY exactly as the interface lists them. get_supplementary_data calls both helpers; its external shape gains only additive keys.

  • Step 4: Full suite → all pass (existing test_school_details.py etc. must not break; if a fixture enumerates yearly-data columns, extend it with the new NaN columns as needed).

  • Step 5: Commit: feat(api): expose progress CIs, KS4 banding/gaps, admissions detail, report-card labels


Task 5: Computed benchmarks helper (TDD)

Files:

  • Modify: backend/data_loader.py (new function)
  • Test: backend/tests/test_benchmarks.py

Interfaces:

  • Produces for Task 6: compute_benchmarks(df) -> dict — pure function over the main dataframe (latest year, state schools), shape:
{
  "source": "state-school average (computed from our dataset)",
  "year": 202425,
  "primary": {
      "disadvantaged_rwm_expected_pct": 46.1,   # weighted by eligible_pupils
      "eal_pct": 22.3,                           # median
      "sen_support_pct": 14.0,                   # median
      "disadvantaged_pct": 24.8,                 # median (FSM6 proxy)
      "median_pupils": 281,                      # median school size
  },
  "secondary": { "median_pupils": 1024, "eal_pct": ..., "sen_support_pct": ..., "disadvantaged_pct": ... },
}
  • Step 1: Failing tests — build a small synthetic df (6 primary rows with known eligible_pupils/rwm_expected_disadvantaged_pct so the weighted average is hand-checkable; a couple of secondary rows flagged by non-null attainment_8_score), assert: weighted disadvantaged average matches hand computation (not the unweighted mean), medians ignore NaN, secondary block lacks the disadvantaged-RWM key, latest-year filtering (rows from an older year must not affect results), and empty df → {}.

  • Step 2: Run → FAIL.

  • Step 3: Implement in data_loader.py:

def compute_benchmarks(df: pd.DataFrame) -> dict:
    """State-school benchmarks computed from our dataset (spec §5/§8.6).
    These are NOT official DfE figures — consumers must label them
    'state-school average (computed from our dataset)'."""
    if df.empty or "year" not in df.columns:
        return {}
    latest_year = df["year"].max()
    d = df[df["year"] == latest_year]
    if d.empty:
        return {}
    is_secondary = d["attainment_8_score"].notna() if "attainment_8_score" in d.columns else pd.Series(False, index=d.index)
    prim, sec = d[~is_secondary], d[is_secondary]

    def _median(sub, col):
        if col not in sub.columns:
            return None
        v = sub[col].median()
        return round(float(v), 1) if pd.notna(v) else None

    def _weighted_disadvantaged(sub):
        if not {"rwm_expected_disadvantaged_pct", "eligible_pupils"} <= set(sub.columns):
            return None
        s = sub.dropna(subset=["rwm_expected_disadvantaged_pct", "eligible_pupils"])
        if s.empty or s["eligible_pupils"].sum() == 0:
            return None
        w = (s["rwm_expected_disadvantaged_pct"] * s["eligible_pupils"]).sum() / s["eligible_pupils"].sum()
        return round(float(w), 1)

    def _block(sub, with_disadvantaged):
        block = {
            "eal_pct": _median(sub, "eal_pct"),
            "sen_support_pct": _median(sub, "sen_support_pct"),
            "disadvantaged_pct": _median(sub, "disadvantaged_pct"),
            "median_pupils": int(sub["total_pupils"].median()) if "total_pupils" in sub.columns and pd.notna(sub["total_pupils"].median()) else None,
        }
        if with_disadvantaged:
            block["disadvantaged_rwm_expected_pct"] = _weighted_disadvantaged(sub)
        return block

    return {
        "source": "state-school average (computed from our dataset)",
        "year": int(latest_year),
        "primary": _block(prim, with_disadvantaged=True),
        "secondary": _block(sec, with_disadvantaged=False),
    }

(Adapt column presence to the real df — sen_support_pct reaches the df via _MAIN_QUERY; confirm and add it there if the KS2 block doesn't already select it, mirroring Task 4's additions.)

  • Step 4: Full suite → pass. Step 5: Commit: feat(api): computed state-school benchmarks

Task 6: Enrich /api/compare + expose GPS/science national averages (TDD)

Files:

  • Modify: backend/app.py (compare_schools ~line 636; get_national_averages ~line 730)
  • Test: backend/tests/test_compare_enrichment.py

Interfaces (response additions, all additive):

  • /api/compare top level gains: "national_averages" (same payload the /api/national-averages endpoint returns — extract the endpoint body into a helper _national_averages_payload(df) and reuse; do not duplicate the logic) and "benchmarks" (Task 5's compute_benchmarks(df)).

  • Each comparison[urn] gains: "ofsted", "census", "admissions", "admissions_history", "deprivation" from get_supplementary_data (one SessionLocal() for the whole request, closed in finally; on exception the five keys are None/[] — mirror the detail endpoint's defensive pattern at app.py:583-590).

  • get_national_averages' KS2 metric list gains "gps_expected_pct", "gps_high_pct", "science_expected_pct" so the England ticks for GPS/science flow once the data exists.

  • Step 1: Failing tests — monkeypatch load_school_data with a two-school primary df (reuse/extend the fixture style of test_school_details.py) and monkeypatch get_supplementary_data to a canned dict; assert on TestClient(app).get("/api/compare?urns=..."):

    • response keeps the existing shape (comparison[urn]["school_info"]["rwm_expected_pct"] etc.),
    • each school gains the five supplementary keys (canned values round-tripped),
    • top-level national_averages and benchmarks present; benchmarks["source"] is the exact provenance string,
    • a supplementary-layer exception (monkeypatched to raise) degrades to ofsted: None etc. with HTTP 200,
    • /api/national-averages includes gps_expected_pct in the primary block when the df/national table provides it (monkeypatch the national-averages source the endpoint reads).
  • Step 2: Run → FAIL. Step 3: Implement per the interfaces. Step 4: Full suite → pass.

  • Step 5: Commit: feat(api): compare endpoint carries supplementary blocks, national averages and benchmarks


Task 7: PR + verification

  • Step 1: Full suite one more time + uv run --with pyyaml python3 -c "import yaml; yaml.safe_load(open('.gitea/workflows/deploy.yml'))" sanity is NOT needed (no workflow changes) — instead run the dbt parse gate again (Task 1 file).
  • Step 2: Push, open PR via the Gitea API (credential-helper basic auth). PR body: the new response shapes (one JSON sketch), the reused-not-duplicated national-averages helper, the provenance rule for benchmarks, deploy note (fields NULL until prod DAGs run post-promotion), and that no e2e change is needed (no user-facing behaviour changes — the compare UI still reads the old fields; the frontend PR carries the journey updates).
  • Step 3: After merge + staging deploy: curl -s https://stx.schoolcompare.co.uk/api/compare?urns=138690,100140 | python3 -m json.tool | head -80 — verify the new keys and that benchmarks.primary.disadvantaged_rwm_expected_pct is plausible (~45-47). Verify /api/national-averages now carries gps_expected_pct/science_expected_pct (values or honest nulls if DfE suppresses them at national level).

Out of scope

  • Frontend rebuild + e2e journeys (next PR — consumes everything this PR exposes).
  • schemas.py METRIC_DEFINITIONS additions for the trends picker (frontend PR decides which of the new columns become picker metrics).
  • CI-based progress banding logic (frontend computes Above/Average/Below from the CI columns; historical years only).