fix(pipeline): decode GIAS extracts as Windows-1252 #182

Merged
tudor merged 1 commits from fix/gias-encoding into main 2026-10-05 06:31:31 +00:00
Owner

Problem

The site shows "St Thomas ŕ Becket" (138950, 149557) where GIAS has "St Thomas à Becket".

GIAS publishes its extracts in Windows-1252 and its download response declares no charset. Both GIAS streams passed io.StringIO(resp.text) to read_csv, so requests guessed the codec (apparent_encoding was windows-1250 on 3 Oct, where byte 0xE0 is "ŕ"). The encoding="latin-1" argument had no effect, because the text was already decoded. The guess is a heuristic and can change between days, garbling other accented names (è→č, ñ→ń, ò→ň …).

Change

  • New tap_uk_gias/gias_csv.py: read_gias_csv(content, logger) decodes the raw bytes as cp1252 and keeps every column a string, with blanks as '' (as before).
  • A byte Windows-1252 leaves undefined becomes U+FFFD and is logged, so one bad name can't fail the daily run and stop every school from refreshing. None of the nine extracts checked (1 Jul to 3 Oct 2026) contains such a byte.
  • Both streams (establishments and links) use it. The module has no singer_sdk import, so CI can test it.

Verification

  • pipeline/tests/test_gias_encoding.py (4 tests): à, ’, é, ° and ç decode correctly, the windows-1250 reading isn't used, values stay strings, and an undefined byte is replaced and logged. All 4 failed before the change and pass after.
  • Ran the installed tap's GIASEstablishmentsStream.get_records on the real 3 Oct extract, with resp.text forced to the windows-1250 reading: 52,586 rows, both St Thomas à Becket schools correct, no "ŕ", no replacement characters.
  • python -m pytest backend/tests pipeline/tests scripts/ci/tests -q: 332 passed.

After merge

The next daily GIAS run rewrites the names. The two schools above should then read "à" on the site.

Independent of #181 (scheduled dbt selectors). The IDACI and FBIT taps read the same extract the same way, apparently only for code columns, so they're left alone here.

🤖 Generated with Claude Code

## Problem The site shows "St Thomas **ŕ** Becket" (138950, 149557) where GIAS has "St Thomas **à** Becket". GIAS publishes its extracts in Windows-1252 and its download response declares no charset. Both GIAS streams passed `io.StringIO(resp.text)` to `read_csv`, so `requests` guessed the codec (`apparent_encoding` was `windows-1250` on 3 Oct, where byte 0xE0 is "ŕ"). The `encoding="latin-1"` argument had no effect, because the text was already decoded. The guess is a heuristic and can change between days, garbling other accented names (è→č, ñ→ń, ò→ň …). ## Change - New `tap_uk_gias/gias_csv.py`: `read_gias_csv(content, logger)` decodes the raw bytes as cp1252 and keeps every column a string, with blanks as `''` (as before). - A byte Windows-1252 leaves undefined becomes U+FFFD and is logged, so one bad name can't fail the daily run and stop every school from refreshing. None of the nine extracts checked (1 Jul to 3 Oct 2026) contains such a byte. - Both streams (establishments and links) use it. The module has no `singer_sdk` import, so CI can test it. ## Verification - `pipeline/tests/test_gias_encoding.py` (4 tests): à, ’, é, ° and ç decode correctly, the windows-1250 reading isn't used, values stay strings, and an undefined byte is replaced and logged. All 4 failed before the change and pass after. - Ran the installed tap's `GIASEstablishmentsStream.get_records` on the real 3 Oct extract, with `resp.text` forced to the windows-1250 reading: 52,586 rows, both St Thomas à Becket schools correct, no "ŕ", no replacement characters. - `python -m pytest backend/tests pipeline/tests scripts/ci/tests -q`: 332 passed. ## After merge The next daily GIAS run rewrites the names. The two schools above should then read "à" on the site. Independent of #181 (scheduled dbt selectors). The IDACI and FBIT taps read the same extract the same way, apparently only for code columns, so they're left alone here. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
tudor added 1 commit 2026-10-04 09:09:52 +00:00
fix(pipeline): decode GIAS extracts as Windows-1252
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m16s
PR Checks / Backend Smoke (pull_request) Successful in 10s
PR Checks / Build Backend (no push) (pull_request) Successful in 18s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m27s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 1m17s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 19s
94bfac9caf
GIAS publishes its CSVs in Windows-1252 and sends no charset. The tap read
resp.text, so requests guessed the codec, and the encoding="latin-1" passed
to read_csv did nothing on already-decoded text. On 3 Oct 2026 the guess was
windows-1250, and "St Thomas à Becket" (138950, 149557) was stored as
"St Thomas ŕ Becket". A different guess on another day would garble other
accented names.

Both streams now decode the downloaded bytes themselves (gias_csv.py). A byte
Windows-1252 leaves undefined becomes U+FFFD with a logged warning instead of
failing the load, so one odd name cannot stop the daily refresh. None of the
nine extracts checked (1 Jul to 3 Oct 2026) contains such a byte.

Checked by running the tap on the real 3 Oct extract with .text forced to
windows-1250: all 52,586 rows decode, with no "ŕ" and no replacement
characters.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

🤖 AI Code Review (Claude Code)

The tap now decodes GIAS CSV bytes explicitly as Windows-1252 through a shared read_gias_csv helper, used by both the establishments and links streams. This replaces resp.text, where requests guessed the codec and turned "à" into "ŕ". Undefined bytes become U+FFFD and are logged instead of failing the load. The change is small and correct, and the tests cover the real byte sequences and the undefined-byte case. I found no blocking issues.

🟡 Minor

  • pipeline/plugins/extractors/tap-uk-gias/tap_uk_gias/gias_csv.py: The codec is now hard-coded to cp1252. If GIAS ever switches to UTF-8 or adds a BOM, the output will be silently mojibaked (for example "Ã" sequences, or "" prefixed to the first header). The U+FFFD warning would not fire, because valid UTF-8 bytes decode cleanly as cp1252. A cheap guard would be to try a strict UTF-8 decode first, or to check for a BOM and warn.
## 🤖 AI Code Review (Claude Code) The tap now decodes GIAS CSV bytes explicitly as Windows-1252 through a shared `read_gias_csv` helper, used by both the establishments and links streams. This replaces `resp.text`, where requests guessed the codec and turned "à" into "ŕ". Undefined bytes become U+FFFD and are logged instead of failing the load. The change is small and correct, and the tests cover the real byte sequences and the undefined-byte case. I found no blocking issues. ### 🟡 Minor - **pipeline/plugins/extractors/tap-uk-gias/tap_uk_gias/gias_csv.py**: The codec is now hard-coded to cp1252. If GIAS ever switches to UTF-8 or adds a BOM, the output will be silently mojibaked (for example "Ã" sequences, or "" prefixed to the first header). The U+FFFD warning would not fire, because valid UTF-8 bytes decode cleanly as cp1252. A cheap guard would be to try a strict UTF-8 decode first, or to check for a BOM and warn.
tudor merged commit b6e48c4930 into main 2026-10-05 06:31:31 +00:00
Sign in to join this conversation.
No Reviewers
No labels
2 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: tudor/school_compare#182