Files
school_compare/docs/superpowers/plans/2026-07-12-compare-data-foundation.md
T

30 KiB

Compare-Screen Data Foundation (Pipeline PR) Implementation Plan

For agentic workers: REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (- [ ]) syntax for tracking.

Goal: Land every pipeline/dbt change the compare-screen redesign needs (spec §5 + §8 of docs/superpowers/specs/2026-07-11-compare-screen-redesign-design.md): promote raw-but-unstored fields to marts, close the national-averages gaps, and wire the Ofsted report-card columns.

Architecture: Meltano Singer taps load raw.* tables; dbt builds stagingmarts (read-only for the backend). All changes here are additive columns/rows — no breaking changes to existing marts. The full dbt build runs on the server via the Airflow DAGs; locally we gate with dbt parse (no DB needed) plus network-only diagnostic scripts.

Tech Stack: Python (Singer SDK taps), dbt-postgres ~1.10 (invoked as python -m dbt.cli.main), Meltano, PostgreSQL.

Global Constraints

  • No new external sources (spec §5): only fields already in the raw schema or in files the taps already download. The one sanctioned tap change is the Ofsted MI report-card columns (spec §5, §8.4) and the legacy-KS2 year addition (same DfE performance-tables source).
  • Additive only: never rename or drop existing mart columns; the backend maps them 1:1 in backend/models.py.
  • Never push to main. Branch: feat/compare-data-foundation; PR checks must pass.
  • Backend models.py changes belong to the follow-up backend PR, not this one.
  • dbt invocation is always python -m dbt.cli.main (a bare dbt resolves to the wrong binary — see pipeline/dags/school_data_pipeline.py:27).
  • EES suppression codes z/c/x must go through the safe_numeric macro.
  • Computed benchmarks (FSM/EAL/SEN medians, disadvantaged national average) are backend work (spec §5) — explicitly out of scope here.

Task 0: Create the branch

Files: none

  • Step 1: git checkout main && git pull && git checkout -b feat/compare-data-foundation

Task 1: Diagnostics — pin the three unknowns

The spec flags three facts we must confirm from the actual files before wiring code: (a) why gps_expected_pct/science_expected_pct are NULL in marts.fact_ks2_national_averages despite being mapped end-to-end; (b) what the KS2 attainment long file calls its subjects/years for 2021/22 and 2022/23 (subject-level 2022/23 is NULL in prod; school-level 2021/22 is absent); (c) the exact report-card column headers in the current Ofsted MI CSV.

Files:

  • Create: pipeline/scripts/diagnose_compare_gaps.py

Interfaces:

  • Produces: a printed findings report; Tasks 5, 6, 7 consume the confirmed column/label names. Precedent: pipeline/scripts/diagnose_ees_ks4.py.

  • Step 1: Write the diagnostic script

"""Diagnose the three data gaps blocking the compare-screen redesign.

Run from repo root (network access required, no DB needed):
    python pipeline/scripts/diagnose_compare_gaps.py
"""
import io
import re
import sys
import zipfile

import pandas as pd
import requests

sys.path.insert(0, "pipeline/plugins/extractors/tap-uk-ees")
sys.path.insert(0, "pipeline/plugins/extractors/tap-uk-ofsted")
from tap_uk_ees.tap import (  # noqa: E402
    _KS2_NATIONAL_COL_MAP,
    _KS2_NATIONAL_CSV_URL,
    download_release_zip,
    get_all_releases,
)
from tap_uk_ofsted.tap import discover_csv_url  # noqa: E402

TIMEOUT = 120


def check_national_gps_science():
    print("\n=== (a) National catalogue CSV: GPS/science columns ===")
    resp = requests.get(_KS2_NATIONAL_CSV_URL, timeout=TIMEOUT)
    resp.raise_for_status()
    df = pd.read_csv(io.BytesIO(resp.content), dtype=str, keep_default_na=False)
    df.columns = [c.strip().lower() for c in df.columns]
    for csv_col in ("pt_gps_exp", "pt_scita_exp", "avg_readscore", "avg_matscore", "avg_gpsscore"):
        status = "PRESENT" if csv_col in df.columns else "MISSING"
        print(f"  {csv_col}: {status}")
    gps_like = [c for c in df.columns if "gps" in c or "scita" in c or "sci" in c]
    print(f"  all gps/science-ish columns: {gps_like}")
    nat = df[df.get("geographic_level", "").str.strip().str.lower() == "national"]
    print(f"  national rows time_periods: {sorted(nat['time_period'].unique())}")
    # Sample the values our map would read for the latest year
    latest = nat[nat["time_period"] == nat["time_period"].max()]
    for csv_col, field in _KS2_NATIONAL_COL_MAP.items():
        val = latest.iloc[0].get(csv_col, "<col missing>") if len(latest) else "<no row>"
        print(f"  {field} <- {csv_col} = {val!r}")


def check_ks2_attainment_years_subjects():
    print("\n=== (b) EES KS2 attainment: years & subject labels ===")
    releases = get_all_releases("key-stage-2-attainment")
    print(f"  releases found: {[r['time_period'] for r in releases]}")
    for release in releases:
        zf = download_release_zip(release["id"])
        name = next((n for n in zf.namelist()
                     if "ks2_school_attainment_data" in n and n.endswith(".csv")), None)
        if not name:
            print(f"  {release['time_period']}: NO school attainment CSV in ZIP")
            continue
        with zf.open(name) as f:
            df = pd.read_csv(f, dtype=str, keep_default_na=False, nrows=200000)
        years = sorted(df["time_period"].unique())
        subjects = sorted(df["subject"].unique())
        print(f"  release {release['time_period']}: time_periods={years}")
        print(f"    subjects={subjects}")


def check_ofsted_report_card_columns():
    print("\n=== (c) Ofsted MI CSV: report-card columns ===")
    url = discover_csv_url()
    print(f"  MI file: {url}")
    resp = requests.get(url, timeout=TIMEOUT)
    resp.raise_for_status()
    df = pd.read_csv(io.BytesIO(resp.content), dtype=str, keep_default_na=False, nrows=5)
    rc_like = [c for c in df.columns
               if re.search(r"report card|inclusion|curriculum|achievement|safeguard|well.?being|governance", c, re.I)]
    print(f"  candidate report-card columns ({len(rc_like)}):")
    for c in rc_like:
        print(f"    - {c!r}")


if __name__ == "__main__":
    check_national_gps_science()
    check_ks2_attainment_years_subjects()
    check_ofsted_report_card_columns()

Note: if _KS2_NATIONAL_CSV_URL is named differently in tap_uk_ees/tap.py (it is defined near the _KS2_NATIONAL_COL_MAP around line ~490), import whatever constant holds the catalogue CSV URL.

  • Step 2: Run it and record findings

Run: python pipeline/scripts/diagnose_compare_gaps.py 2>&1 | tee /tmp/compare-gaps-findings.txt Expected: three sections printed. Paste the findings as a comment block at the bottom of the script (so they're committed evidence), e.g. # FINDINGS 2026-07-12: pt_gps_exp MISSING (actual col: ...), 202122 present in release X, rc columns: [...].

  • Step 3: Commit
git add pipeline/scripts/diagnose_compare_gaps.py
git commit -m "chore(pipeline): diagnostic for compare-screen data gaps"

Task 2: Admissions preference detail → mart

Staging already extracts second_preference_offers, third_preference_offers, total_offers (stg_ees_admissions.sql:26-29) — the mart drops them. The cross-LA fields are declared in the tap (all_applications_from_another_LA, offers_to_applicants_from_another_LA) but not selected in staging.

Files:

  • Modify: pipeline/transform/models/staging/stg_ees_admissions.sql (after line 33, in renamed)
  • Modify: pipeline/transform/models/marts/fact_admissions.sql
  • Modify: pipeline/transform/models/marts/_marts_schema.yml (fact_admissions block, ~line 120)

Interfaces:

  • Produces mart columns: total_offers int, second_preference_offers int, third_preference_offers int, cross_la_applications int, cross_la_offers int. The backend PR will map these in FactAdmissions.

  • Step 1: Add cross-LA columns to staging

In stg_ees_admissions.sql, after the first_preference_applications line (line 33):

        -- Cross-borough demand: applications naming this school from families
        -- living in another local authority, and offers made to them.
        {{ safe_numeric('"all_applications_from_another_LA"') }}::integer as cross_la_applications,
        {{ safe_numeric('"offers_to_applicants_from_another_LA"') }}::integer as cross_la_offers,

(Quote the identifiers — the tap emits them with mixed case, same trap as FSM_eligible_percent, see the header comment in that file. If dbt parse or the DAG run later shows the raw columns are lower-cased in Postgres, drop the double quotes.)

  • Step 2: Pass everything through the mart

Replace the full select list in fact_admissions.sql:

-- Mart: School admissions — one row per URN per year

select
    urn,
    year,
    school_phase,
    places_offered,
    total_offers,
    total_applications,
    first_preference_applications,
    first_preference_offers,
    second_preference_offers,
    third_preference_offers,
    cross_la_applications,
    cross_la_offers,
    first_preference_offer_pct,
    oversubscription_ratio,
    oversubscribed,
    admissions_policy
from {{ ref('stg_ees_admissions') }}
  • Step 3: Add schema tests

In _marts_schema.yml under fact_admissions.columns, append:

      - name: second_preference_offers
      - name: third_preference_offers
      - name: cross_la_applications
      - name: cross_la_offers
      - name: total_offers
  • Step 4: Parse gate

Run: cd pipeline/transform && python -m dbt.cli.main parse --profiles-dir . Expected: Done. with no compilation errors.

  • Step 5: Commit
git add pipeline/transform/models/staging/stg_ees_admissions.sql pipeline/transform/models/marts/fact_admissions.sql pipeline/transform/models/marts/_marts_schema.yml
git commit -m "feat(pipeline): admissions preference breakdown and cross-LA demand in marts"

Task 3: KS2 progress confidence intervals + writing working-towards

The tap already emits progress_measure_lower_conf_interval, progress_measure_upper_conf_interval, working_towards_expected_standard_pupil_percent (tap.py:203-206). The staging pivot drops them. These power the CI-based Above/Average/Below progress chips (spec §8, first-review item on statistical honesty).

Files:

  • Modify: pipeline/transform/models/staging/stg_ees_ks2.sql (inside the pivoted CTE, next to each subject's progress_measure_score case, lines ~41/55/72, and in the final select ~lines 145-152)
  • Modify: pipeline/transform/models/marts/fact_ks2_performance.sql
  • Modify: pipeline/transform/models/marts/_marts_schema.yml (fact_ks2_performance block, ~line 82)

Interfaces:

  • Produces mart columns: reading_progress_lower_ci, reading_progress_upper_ci, writing_progress_lower_ci, writing_progress_upper_ci, maths_progress_lower_ci, maths_progress_upper_ci (float), writing_working_towards_pct (float).

  • Step 1: Add pivot cases in staging

After the reading_progress case (line ~41), add:

        max(case when subject = 'Reading'
                  and breakdown_topic = 'All pupils' and breakdown = 'Total'
            then {{ safe_numeric('progress_measure_lower_conf_interval') }} end) as reading_progress_lower_ci,
        max(case when subject = 'Reading'
                  and breakdown_topic = 'All pupils' and breakdown = 'Total'
            then {{ safe_numeric('progress_measure_upper_conf_interval') }} end) as reading_progress_upper_ci,

After the writing_progress case (line ~55), add:

        max(case when subject = 'Writing'
                  and breakdown_topic = 'All pupils' and breakdown = 'Total'
            then {{ safe_numeric('progress_measure_lower_conf_interval') }} end) as writing_progress_lower_ci,
        max(case when subject = 'Writing'
                  and breakdown_topic = 'All pupils' and breakdown = 'Total'
            then {{ safe_numeric('progress_measure_upper_conf_interval') }} end) as writing_progress_upper_ci,
        max(case when subject = 'Writing'
                  and breakdown_topic = 'All pupils' and breakdown = 'Total'
            then {{ safe_numeric('working_towards_expected_standard_pupil_percent') }} end) as writing_working_towards_pct,

After the maths_progress case (line ~72), add:

        max(case when subject = 'Maths'
                  and breakdown_topic = 'All pupils' and breakdown = 'Total'
            then {{ safe_numeric('progress_measure_lower_conf_interval') }} end) as maths_progress_lower_ci,
        max(case when subject = 'Maths'
                  and breakdown_topic = 'All pupils' and breakdown = 'Total'
            then {{ safe_numeric('progress_measure_upper_conf_interval') }} end) as maths_progress_upper_ci,

Then add the seven new columns to the model's final select (next to the existing p.reading_progress / p.writing_progress / p.maths_progress lines ~145-152):

    p.reading_progress_lower_ci,
    p.reading_progress_upper_ci,
    p.writing_progress_lower_ci,
    p.writing_progress_upper_ci,
    p.writing_working_towards_pct,
    p.maths_progress_lower_ci,
    p.maths_progress_upper_ci,
  • Step 2: Pass through the mart

In fact_ks2_performance.sql, add the same seven column names to the select list immediately after the existing maths_progress line (this mart selects staging columns by name; match the file's existing alias style — if columns are selected bare, add them bare).

  • Step 3: Schema tests

In _marts_schema.yml under fact_ks2_performance.columns, append the seven names (no tests beyond presence — values are legitimately NULL for 2023/24+ since progress measures ended with 2022/23, spec §4.3):

      - name: reading_progress_lower_ci
      - name: reading_progress_upper_ci
      - name: writing_progress_lower_ci
      - name: writing_progress_upper_ci
      - name: writing_working_towards_pct
      - name: maths_progress_lower_ci
      - name: maths_progress_upper_ci
  • Step 4: Parse gate

Run: cd pipeline/transform && python -m dbt.cli.main parse --profiles-dir . Expected: Done.

  • Step 5: Commit
git add pipeline/transform/models/staging/stg_ees_ks2.sql pipeline/transform/models/marts/fact_ks2_performance.sql pipeline/transform/models/marts/_marts_schema.yml
git commit -m "feat(pipeline): KS2 progress confidence intervals and writing working-towards"

Task 4: KS4 — Progress 8 banding and disadvantage gaps

The tap's ees_ks4_info stream already declares progress8_banding (DfE's own "well above average … well below average" label — the ready-made secondary chip), attainment8_diffn and progress8_diffn (tap.py:338-340). Wire them through staging into the mart.

Files:

  • Modify: pipeline/transform/models/staging/stg_ees_ks4.sql (the CTE that reads ees_ks4_info — the same one that already surfaces sen_pct; add three columns to its select and to the final joined select)
  • Modify: pipeline/transform/models/marts/fact_ks4_performance.sql (add after progress_8_upper_ci)
  • Modify: pipeline/transform/models/marts/_marts_schema.yml (fact_ks4_performance block, ~line 93)

Interfaces:

  • Produces mart columns: progress_8_banding text, attainment_8_disadvantage_gap float, progress_8_disadvantage_gap float.

  • Step 1: Staging — select from the info source

In the info CTE of stg_ees_ks4.sql add:

        nullif(trim(progress8_banding), '') as progress_8_banding,
        {{ safe_numeric('attainment8_diffn') }} as attainment_8_disadvantage_gap,
        {{ safe_numeric('progress8_diffn') }} as progress_8_disadvantage_gap,

and add the three names to the model's final select (aliased the same way the CTE's other columns are).

  • Step 2: Mart passthrough

In fact_ks4_performance.sql, after the progress_8_upper_ci, line:

    progress_8_banding,
    attainment_8_disadvantage_gap,
    progress_8_disadvantage_gap,
  • Step 3: Schema tests — append the three names under fact_ks4_performance.columns, plus an accepted-values guard that tolerates NULL:
      - name: progress_8_banding
        tests:
          - accepted_values:
              values: ['Well above average', 'Above average', 'Average', 'Below average', 'Well below average']
              config:
                where: "progress_8_banding is not null"
      - name: attainment_8_disadvantage_gap
      - name: progress_8_disadvantage_gap

(If the DAG run later shows different capitalisation in the data, fix the accepted values to match the data, not vice versa.)

  • Step 4: Parse gatecd pipeline/transform && python -m dbt.cli.main parse --profiles-dir .Done.

  • Step 5: Commit

git add pipeline/transform/models/staging/stg_ees_ks4.sql pipeline/transform/models/marts/fact_ks4_performance.sql pipeline/transform/models/marts/_marts_schema.yml
git commit -m "feat(pipeline): Progress 8 banding and KS4 disadvantage gaps in marts"

Task 5: National averages — 2015/16 row and GPS/science/scaled-score fix

Two changes. (1) stg_ees_ks2_national.sql:34 filters >= 201617, which is exactly why the England line starts a year late (2015/16 RWM = 53% exists in the catalogue). (2) GPS/science expected are NULL in prod despite full end-to-end mapping — Task 1's findings say whether the catalogue CSV column names differ from _KS2_NATIONAL_COL_MAP (pt_gps_exp, pt_scita_exp) or whether values are suppressed at source.

Files:

  • Modify: pipeline/transform/models/staging/stg_ees_ks2_national.sql:34
  • Modify (conditional on Task 1 findings): pipeline/plugins/extractors/tap-uk-ees/tap_uk_ees/tap.py (_KS2_NATIONAL_COL_MAP)

Interfaces:

  • Produces: a 201516 row in marts.fact_ks2_national_averages; non-NULL gps_expected_pct, science_expected_pct, reading_avg_score, maths_avg_score, gps_avg_score for years the DfE publishes them. Backend/frontend consume via /api/national-averages unchanged (additive year + newly non-NULL fields).

  • Step 1: Widen the year filter

In stg_ees_ks2_national.sql, change line 34:

  and cast(trim(time_period) as integer) >= 201516

(2015/16 was the first year of the current expected-standard tests; nothing earlier is comparable, so keep a floor.)

  • Step 2: Fix the column map per Task 1 findings

If Task 1 reported the actual CSV column names for GPS/science/scaled scores differ, update _KS2_NATIONAL_COL_MAP in tap.py accordingly, e.g. (illustrative — use the diagnosed names):

_KS2_NATIONAL_COL_MAP = {
    # ... existing entries ...
    "pt_gps_exp": "gps_expected_pct",      # replace key with diagnosed name
    "pt_scita_exp": "science_expected_pct",  # replace key with diagnosed name
}

If Task 1 showed the columns are present but suppressed (x) at national level for all years, instead delete the two entries from the map, delete the corresponding lines from stg_ees_ks2_national.sql and fact_ks2_national_averages.sql, and record in the PR description that GPS/science England ticks stay "not in dataset" (the mockups already carry that caveat).

  • Step 3: Parse gatecd pipeline/transform && python -m dbt.cli.main parse --profiles-dir .Done.

  • Step 4: Commit

git add pipeline/transform/models/staging/stg_ees_ks2_national.sql pipeline/plugins/extractors/tap-uk-ees/tap_uk_ees/tap.py
git commit -m "fix(pipeline): include 2015/16 national averages; fix GPS/science national mapping"

Task 6: Legacy KS2 — load the 2021/22 school-level year

School-level 2021/22 exists in DfE performance-tables archives (same source as the four legacy years already loaded) but in neither our legacy config (stops at 201819, pipeline/meltano.yml:33-37) nor EES (starts 2022/23) — unless Task 1's finding (b) showed an EES release carrying 202122, in which case skip this task and note why in the PR.

The legacy URLs point at the self-hosted filebrowser (10.0.1.224:8081) — the 2021/22 DfE archive must be uploaded there first; this is the one human dependency in this plan.

Files:

  • Modify: pipeline/meltano.yml (legacy_ks2_urls block, line ~33)

Interfaces:

  • Produces: raw.legacy_ks2 rows with year = '202122', flowing through stg_legacy_ks2fact_ks2_performance unchanged (the stream maps old column names already; 2021/22 CSVs use the same PTRWM_EXP-style headers as 2018/19).

  • Step 1: Verify the 2021/22 CSV headers match _LEGACY_KS2_COLUMN_MAP

Download the DfE 2021/22 KS2 revised archive (gov.uk "Compare School Performance data download": 2021-2022 all-schools ZIP), then:

Run: python -c "import zipfile,io,pandas as pd; zf=zipfile.ZipFile('/path/to/2021-2022.zip'); n=[x for x in zf.namelist() if 'ks2final' in x.lower() and x.endswith('.csv')][0]; df=pd.read_csv(zf.open(n), dtype=str, nrows=5); import sys; sys.path.insert(0,'pipeline/plugins/extractors/tap-uk-ees'); from tap_uk_ees.tap import _LEGACY_KS2_COLUMN_MAP as m; missing=[c for c in m if c not in df.columns]; print('missing legacy columns:', missing)" Expected: missing legacy columns: [] (progress columns READPROG etc. may legitimately be missing/blank in 2021/22 — acceptable, they load as NULL).

  • Step 2: Upload the archive to the filebrowser and add the config entry

In pipeline/meltano.yml under legacy_ks2_urls, add (with the real share URL from the filebrowser upload):

          "202122": "http://10.0.1.224:8081/filebrowser/api/public/dl/<SHARE_ID>?inline=true"
  • Step 3: Commit
git add pipeline/meltano.yml
git commit -m "feat(pipeline): load 2021/22 school-level KS2 from legacy performance tables"
  • Step 4 (only if Task 1(b) showed 2022/23 subject labels differ): widen the subject matchers in stg_ees_ks2.sql the same way GPS already is (subject ilike '%grammar%' or subject = 'GPS'), e.g. subject in ('Reading', 'reading') → use the diagnosed labels. Parse-gate and commit as fix(pipeline): match 2022/23 KS2 subject labels.

Task 7: Ofsted report-card columns (rc_*)

Resolves the tap TODO (stg_ofsted_inspections.sql:37). The marts/backed columns already exist as stubs; this wires real values. Uses Task 1(c)'s confirmed MI column names — the candidates below follow the MI file's existing naming style and must be corrected against the diagnostic output.

Files:

  • Modify: pipeline/plugins/extractors/tap-uk-ofsted/tap_uk_ofsted/tap.py (COLUMN_PRIORITY ~line 19-72, schema ~line 100-114)
  • Create: pipeline/transform/macros/parse_report_card_grade.sql
  • Modify: pipeline/transform/models/staging/stg_ofsted_inspections.sql:36-46

Interfaces:

  • Produces mart columns (already declared in fact_ofsted_inspection): rc_safeguarding_met boolean, and rc_inclusionrc_sixth_form as integers on the 5-point scale 1=Exceptional, 2=Strong standard, 3=Expected standard, 4=Needs attention/Attention needed, 5=Urgent improvement. The backend translates codes to labels (same pattern as gias_codes.py), verifying wording against Ofsted's published toolkit (spec §8.4).

  • Step 1: Add tap column mappings

In COLUMN_PRIORITY add (replace candidate strings with Task 1(c)'s exact headers — keep them as priority lists so older files degrade to blank):

    "rc_safeguarding_met": ["Report card safeguarding", "Safeguarding"],
    "rc_inclusion": ["Report card inclusion", "Inclusion"],
    "rc_curriculum_teaching": ["Report card curriculum and teaching", "Curriculum and teaching"],
    "rc_achievement": ["Report card achievement", "Achievement"],
    "rc_attendance_behaviour": ["Report card attendance and behaviour", "Attendance and behaviour"],
    "rc_personal_development": ["Report card personal development and well-being", "Personal development and well-being"],
    "rc_leadership_governance": ["Report card leadership and governance", "Leadership and governance"],
    "rc_early_years": ["Report card early years", "Early years"],
    "rc_sixth_form": ["Report card sixth form", "Sixth form"],

And in the stream schema (next to report_url, ~line 114):

        th.Property("rc_safeguarding_met", th.StringType),
        th.Property("rc_inclusion", th.StringType),
        th.Property("rc_curriculum_teaching", th.StringType),
        th.Property("rc_achievement", th.StringType),
        th.Property("rc_attendance_behaviour", th.StringType),
        th.Property("rc_personal_development", th.StringType),
        th.Property("rc_leadership_governance", th.StringType),
        th.Property("rc_early_years", th.StringType),
        th.Property("rc_sixth_form", th.StringType),
  • Step 2: Write the grade-parsing macro

pipeline/transform/macros/parse_report_card_grade.sql:

{% macro parse_report_card_grade(column_name) %}
    case lower(trim(nullif({{ column_name }}, 'NULL')))
        when 'exceptional' then 1
        when 'strong standard' then 2
        when 'expected standard' then 3
        when 'needs attention' then 4
        when 'attention needed' then 4
        when 'urgent improvement' then 5
    end
{% endmacro %}
  • Step 3: Wire staging

Replace stg_ofsted_inspections.sql lines 36-46 (the NULL stubs) with:

        -- Report Card fields (post-Nov 2025 framework), 5-point scale:
        -- 1 Exceptional · 2 Strong standard · 3 Expected standard
        -- · 4 Needs attention · 5 Urgent improvement
        (lower(trim(nullif(rc_safeguarding_met, 'NULL'))) = 'met')        as rc_safeguarding_met,
        {{ parse_report_card_grade('rc_inclusion') }}::integer            as rc_inclusion,
        {{ parse_report_card_grade('rc_curriculum_teaching') }}::integer  as rc_curriculum_teaching,
        {{ parse_report_card_grade('rc_achievement') }}::integer          as rc_achievement,
        {{ parse_report_card_grade('rc_attendance_behaviour') }}::integer as rc_attendance_behaviour,
        {{ parse_report_card_grade('rc_personal_development') }}::integer as rc_personal_development,
        {{ parse_report_card_grade('rc_leadership_governance') }}::integer as rc_leadership_governance,
        {{ parse_report_card_grade('rc_early_years') }}::integer          as rc_early_years,
        {{ parse_report_card_grade('rc_sixth_form') }}::integer           as rc_sixth_form,

Note rc_safeguarding_met becomes boolean (NULL when blank) — matching fact_ofsted_inspection's rc_safeguarding_met Boolean column. If fact_ofsted_inspection.sql casts these columns, align its casts too (inspect that model; it currently passes the text stubs through).

  • Step 4: Parse gate + tap smoke test

Run: cd pipeline/transform && python -m dbt.cli.main parse --profiles-dir .Done. Run: python -c "import sys; sys.path.insert(0,'pipeline/plugins/extractors/tap-uk-ofsted'); from tap_uk_ofsted.tap import COLUMN_PRIORITY; assert 'rc_inclusion' in COLUMN_PRIORITY; print('ok')"ok

  • Step 5: Commit
git add pipeline/plugins/extractors/tap-uk-ofsted/tap_uk_ofsted/tap.py pipeline/transform/macros/parse_report_card_grade.sql pipeline/transform/models/staging/stg_ofsted_inspections.sql
git commit -m "feat(pipeline): extract Ofsted report-card judgements (rc_* columns)"

Task 8: PR + post-merge verification

Files: none new

  • Step 1: Push and open the PR (Gitea — use the git credential helper + basic-auth API pattern; token-header auth 401s):
git push -u origin feat/compare-data-foundation
# then create the PR via the Gitea API with basic auth from `git credential fill`

PR body: link spec §5/§8, list the new mart columns, note the Task 6 human dependency (filebrowser upload) and the Task 1 findings file.

  • Step 2: After merge, verify the DAG run picked everything up

The daily/monthly DAGs rebuild the affected models (pipeline/dags/school_data_pipeline.py). Spot-check via the public API (production after promotion, staging first at stx.schoolcompare.co.uk — note external /api is broken at the staging proxy, so check staging from the host):

# 2015/16 national row exists
curl -sL "https://www.schoolcompare.co.uk/api/national-averages" | python3 -c "import json,sys; d=json.load(sys.stdin); assert any(r['year']==201516 and r['primary'] for r in d['by_year']), '2015/16 missing'; print('201516 ok')"
# 2021/22 school rows exist (Barclay)
curl -sL "https://www.schoolcompare.co.uk/api/schools/138690" | python3 -c "import json,sys; d=json.load(sys.stdin); ys=[r['year'] for r in d['yearly_data']]; assert 202122 in [int(y) for y in ys], ys; print('202122 ok')"

(The admissions/CI/KS4/rc_* columns aren't API-visible until the backend PR maps them — verify those directly in Postgres from the pipeline host: select count(*) from marts.fact_admissions where second_preference_offers is not null; etc.)

  • Step 3: Update the spec — tick off the §5 promotions this PR delivered (edit the spec's promotion list to note "landed in PR #NN") and commit to main via a docs PR or alongside the backend PR.

Out of scope (next plans)

  1. Backend PR: map new columns in backend/models.py, extend /api/compare with supplementary blocks + national_averages, computed benchmarks (FSM/EAL/SEN/size medians, disadvantaged national average), CI-based progress banding, report-card label translation (verify against Ofsted toolkit), Ofsted provider-page URLs, graded-vs-ungraded surfacing.
  2. Frontend PR: rebuild /compare per the mockups + e2e journeys (promotion gate).
  3. Separate bug fix: third school's series not rendering on the current production chart.
  4. Post-v1 (spec): census ethnicity/young-carer promotion, IDACI display, attendance section, gender-split/absence tier-2 measures.
  5. Already in marts, no work needed: KS4 EBacc entry/APS, grade 5+ English & maths, Progress 8 CIs — fact_ks4_performance carries them today; only the backend needs to expose them.