7.9 KiB
GIAS Code Dictionaries — Codes in Marts, Names in Code
Date: 2026-07-09 Status: Approved design
Goal
Six GIAS classification fields are stored in the marts as repeated name strings. Replace them with the official DfE integer codes and translate code → name in application code. After this change the marts carry only codes for:
| GIAS field | Today (marts, string) | After (marts, int) |
|---|---|---|
TypeOfEstablishment (name) |
dim_school.school_type |
school_type_code |
EstablishmentStatus (name) |
dim_school.status |
status_code |
PhaseOfEducation (name) |
dim_school.phase |
phase_code |
OfficialSixthForm (name) |
(already reduced to has_sixth_form bool) |
official_sixth_form_code in staging only; mart keeps the bool |
ReligiousCharacter (name) |
dim_school.religious_character |
religious_character_code |
AdmissionsPolicy (name) |
dim_school.admissions_policy |
admissions_policy_code |
Motivation: smaller marts and stable enum values for filtering. (Honest sizing note: at ~25k open schools the raw performance win is modest; the durable benefits are storage, DfE-governed vocabulary, and filter values that can't drift with GIAS renames.)
Decisions (made during brainstorming)
- GIAS native codes, not custom enums. The GIAS bulk CSV publishes an
official
X (code)column beside everyX (name)column. We ingest the DfE's own codes; no invented mapping to maintain. - Translation lives in the backend at the API boundary. The API keeps serving today's name strings; the frontend, e2e journeys, and API consumers are untouched.
Design
1. Tap (Singer schema)
Add the six (code) columns to GIASEstablishmentsStream.schema in
pipeline/plugins/extractors/tap-uk-gias/tap_uk_gias/tap.py:
"TypeOfEstablishment (code)", "EstablishmentStatus (code)",
"PhaseOfEducation (code)", "OfficialSixthForm (code)",
"ReligiousCharacter (code)", "AdmissionsPolicy (code)"
The (name) columns stay declared — raw keeps both so we can detect
dictionary drift (§4) and regenerate dictionaries from live data.
2. Staging (stg_gias_establishments.sql)
- Add int casts:
school_type_code,status_code,phase_code,official_sixth_form_code,religious_character_code,admissions_policy_code(allcast(nullif(trim(...), '') as integer)). - Remove the corresponding name columns from the staging select
(
school_type,status,phase,official_sixth_form,religious_character,admissions_policy). Names live only in raw.
3. Marts
dim_school stores codes only:
school_type_code,status_code,phase_code,religious_character_code,admissions_policy_codereplace their string columns.- Status filter becomes
where status_code in (<open>, <proposed-to-close>). The numeric values are read from live raw data at implementation time (select distinct "EstablishmentStatus (code)", "EstablishmentStatus (name)"), never assumed from memory. Same filter indim_location. has_sixth_formderives fromofficial_sixth_form_code(<has-code>→ true,<does-not>/<not-applicable>→ false, null →statutory_high_age >= 18fallback). Thelower(trim(...))string guard becomes obsolete and is removed.phase_codederivation keeps today's cascade but emits codes:- GIAS
phase_codewhen it is a real value (not the not-applicable code); - statutory-age inference emits the matching GIAS code (Primary / Secondary / All-through — numeric values confirmed from live data at implementation);
- school-name heuristics (unchanged — they match
school_name, which is not one of the six fields) emit the same codes; - else null.
- GIAS
- dbt schema tests:
accepted_values(severity warn) on every code column, values taken from the dictionary;not_nullwarn onphase_code(mirrors today's phase test);has_sixth_formtests unchanged.
dim_location: only the status filter changes (must stay byte-identical
to dim_school's — the API inner-joins the two).
4. Dictionaries
Canonical module: backend/gias_codes.py
ESTABLISHMENT_STATUS: dict[int, str]
SCHOOL_TYPE: dict[int, str]
PHASE_OF_EDUCATION: dict[int, str]
OFFICIAL_SIXTH_FORM: dict[int, str]
RELIGIOUS_CHARACTER: dict[int, str]
ADMISSIONS_POLICY: dict[int, str]
def translate(code: int | None, mapping: dict[int, str]) -> str | None:
"""None -> None; unknown code -> 'Unknown (<code>)' + warning log."""
- Contents are generated from live raw data
(
SELECT DISTINCT code, name FROM raw.gias_establishments ...per field) and sanity-checked against the DfE GIAS registers. Names must be byte-identical to what the API serves today. - Unknown codes never blank the UI:
translatereturns"Unknown (<code>)"and logs, so a new DfE value degrades gracefully.
Pipeline copy: pipeline/scripts/gias_codes.py
The app and pipeline Docker images have disjoint build contexts
(Dockerfile copies backend/; pipeline/Dockerfile copies pipeline/),
so the Typesense sync cannot import the backend module. It gets a
byte-identical copy, and a backend unit test asserts
backend/gias_codes.py and pipeline/scripts/gias_codes.py have identical
content — drift fails CI. (Deliberately chosen over codegen: six dicts do
not justify build machinery.)
Seed for drift detection: pipeline/transform/seeds/gias_code_names.csv
Columns field,code,name mirroring the dictionary. A dbt test (severity
warn) compares live raw (code, name) pairs against the seed; when DfE adds
or renames a value the nightly run warns, prompting a dictionary + seed
update in one PR.
5. Backend translation (API contract unchanged)
_MAIN_QUERYselects the code columns instead of the name columns.load_school_data_as_dataframe()translates immediately afterpd.read_sql, writing today's column names:
df["phase"] = df["phase_code"].map(...)
df["school_type"] = df["school_type_code"].map(...) # then normalize_school_type as today
df["status"] = df["status_code"].map(...)
df["religious_denomination"] = df["religious_character_code"].map(...)
df["admissions_policy"] = df["admissions_policy_code"].map(...)
Everything downstream — PHASE_GROUPS, filters, payload builders,
/api/filters, frontend, e2e — sees exactly today's strings. No frontend
changes.
backend/models.pyDimSchool: string columns replaced by*_code = Column(Integer).
6. Typesense sync
pipeline/scripts/sync_typesense.py selects phase, school_type,
religious_character today. It switches to the code columns and translates
via pipeline/scripts/gias_codes.py before indexing, so facet values in
search are unchanged.
7. Rollout
- No DB migration: marts are full-rebuild tables.
- Deploy window: until the first post-merge pipeline run, the old marts
still carry string columns while the new backend queries code columns, so
the backend's query fails and it serves empty data (the one-column retry
built for
has_sixth_formdoesn't generalise to six columns, and a full old-schema fallback query isn't worth it). Decision: accept the window and close it operationally — the runbook is merge → deploy → triggerschool_data_dailyimmediately. The DAG's final step already calls/api/admin/reload, so the backend recovers without a restart. - Tests: backend unit tests for
translate()(known / unknown / None), payload tests asserting names still served, the file-parity test, dbt schema/seed tests. Frontend: no changes; existing Jest suite is the regression net.
Out of scope
- Recoding other string columns (
gender,urban_rural,nursery_provision,local_authority_name…) — same pattern can follow later if this proves out. - Collapsing academy subtypes (today's
normalize_school_type) — kept as-is, applied after translation. - Serving codes through the API — the contract deliberately keeps names.