The site shows "St Thomas ŕ Becket" (138950, 149557) where GIAS has "St Thomas à Becket".
GIAS publishes its extracts in Windows-1252 and its download response declares no charset. Both GIAS streams passed io.StringIO(resp.text) to read_csv, so requests guessed the codec (apparent_encoding was windows-1250 on 3 Oct, where byte 0xE0 is "ŕ"). The encoding="latin-1" argument had no effect, because the text was already decoded. The guess is a heuristic and can change between days, garbling other accented names (è→č, ñ→ń, ò→ň …).
Change
New tap_uk_gias/gias_csv.py: read_gias_csv(content, logger) decodes the raw bytes as cp1252 and keeps every column a string, with blanks as '' (as before).
A byte Windows-1252 leaves undefined becomes U+FFFD and is logged, so one bad name can't fail the daily run and stop every school from refreshing. None of the nine extracts checked (1 Jul to 3 Oct 2026) contains such a byte.
Both streams (establishments and links) use it. The module has no singer_sdk import, so CI can test it.
Verification
pipeline/tests/test_gias_encoding.py (4 tests): à, ’, é, ° and ç decode correctly, the windows-1250 reading isn't used, values stay strings, and an undefined byte is replaced and logged. All 4 failed before the change and pass after.
Ran the installed tap's GIASEstablishmentsStream.get_records on the real 3 Oct extract, with resp.text forced to the windows-1250 reading: 52,586 rows, both St Thomas à Becket schools correct, no "ŕ", no replacement characters.
The next daily GIAS run rewrites the names. The two schools above should then read "à" on the site.
Independent of #181 (scheduled dbt selectors). The IDACI and FBIT taps read the same extract the same way, apparently only for code columns, so they're left alone here.
## Problem
The site shows "St Thomas **ŕ** Becket" (138950, 149557) where GIAS has "St Thomas **à** Becket".
GIAS publishes its extracts in Windows-1252 and its download response declares no charset. Both GIAS streams passed `io.StringIO(resp.text)` to `read_csv`, so `requests` guessed the codec (`apparent_encoding` was `windows-1250` on 3 Oct, where byte 0xE0 is "ŕ"). The `encoding="latin-1"` argument had no effect, because the text was already decoded. The guess is a heuristic and can change between days, garbling other accented names (è→č, ñ→ń, ò→ň …).
## Change
- New `tap_uk_gias/gias_csv.py`: `read_gias_csv(content, logger)` decodes the raw bytes as cp1252 and keeps every column a string, with blanks as `''` (as before).
- A byte Windows-1252 leaves undefined becomes U+FFFD and is logged, so one bad name can't fail the daily run and stop every school from refreshing. None of the nine extracts checked (1 Jul to 3 Oct 2026) contains such a byte.
- Both streams (establishments and links) use it. The module has no `singer_sdk` import, so CI can test it.
## Verification
- `pipeline/tests/test_gias_encoding.py` (4 tests): à, ’, é, ° and ç decode correctly, the windows-1250 reading isn't used, values stay strings, and an undefined byte is replaced and logged. All 4 failed before the change and pass after.
- Ran the installed tap's `GIASEstablishmentsStream.get_records` on the real 3 Oct extract, with `resp.text` forced to the windows-1250 reading: 52,586 rows, both St Thomas à Becket schools correct, no "ŕ", no replacement characters.
- `python -m pytest backend/tests pipeline/tests scripts/ci/tests -q`: 332 passed.
## After merge
The next daily GIAS run rewrites the names. The two schools above should then read "à" on the site.
Independent of #181 (scheduled dbt selectors). The IDACI and FBIT taps read the same extract the same way, apparently only for code columns, so they're left alone here.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
GIAS publishes its CSVs in Windows-1252 and sends no charset. The tap read
resp.text, so requests guessed the codec, and the encoding="latin-1" passed
to read_csv did nothing on already-decoded text. On 3 Oct 2026 the guess was
windows-1250, and "St Thomas à Becket" (138950, 149557) was stored as
"St Thomas ŕ Becket". A different guess on another day would garble other
accented names.
Both streams now decode the downloaded bytes themselves (gias_csv.py). A byte
Windows-1252 leaves undefined becomes U+FFFD with a logged warning instead of
failing the load, so one odd name cannot stop the daily refresh. None of the
nine extracts checked (1 Jul to 3 Oct 2026) contains such a byte.
Checked by running the tap on the real 3 Oct extract with .text forced to
windows-1250: all 52,586 rows decode, with no "ŕ" and no replacement
characters.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The tap now decodes GIAS CSV bytes explicitly as Windows-1252 through a shared read_gias_csv helper, used by both the establishments and links streams. This replaces resp.text, where requests guessed the codec and turned "à" into "ŕ". Undefined bytes become U+FFFD and are logged instead of failing the load. The change is small and correct, and the tests cover the real byte sequences and the undefined-byte case. I found no blocking issues.
🟡 Minor
pipeline/plugins/extractors/tap-uk-gias/tap_uk_gias/gias_csv.py: The codec is now hard-coded to cp1252. If GIAS ever switches to UTF-8 or adds a BOM, the output will be silently mojibaked (for example "Ã" sequences, or "" prefixed to the first header). The U+FFFD warning would not fire, because valid UTF-8 bytes decode cleanly as cp1252. A cheap guard would be to try a strict UTF-8 decode first, or to check for a BOM and warn.
## 🤖 AI Code Review (Claude Code)
The tap now decodes GIAS CSV bytes explicitly as Windows-1252 through a shared `read_gias_csv` helper, used by both the establishments and links streams. This replaces `resp.text`, where requests guessed the codec and turned "à" into "ŕ". Undefined bytes become U+FFFD and are logged instead of failing the load. The change is small and correct, and the tests cover the real byte sequences and the undefined-byte case. I found no blocking issues.
### 🟡 Minor
- **pipeline/plugins/extractors/tap-uk-gias/tap_uk_gias/gias_csv.py**: The codec is now hard-coded to cp1252. If GIAS ever switches to UTF-8 or adds a BOM, the output will be silently mojibaked (for example "Ã" sequences, or "" prefixed to the first header). The U+FFFD warning would not fire, because valid UTF-8 bytes decode cleanly as cp1252. A cheap guard would be to try a strict UTF-8 decode first, or to check for a BOM and warn.
tudor
merged commit b6e48c4930 into main2026-10-05 06:31:31 +00:00
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Problem
The site shows "St Thomas ŕ Becket" (138950, 149557) where GIAS has "St Thomas à Becket".
GIAS publishes its extracts in Windows-1252 and its download response declares no charset. Both GIAS streams passed
io.StringIO(resp.text)toread_csv, sorequestsguessed the codec (apparent_encodingwaswindows-1250on 3 Oct, where byte 0xE0 is "ŕ"). Theencoding="latin-1"argument had no effect, because the text was already decoded. The guess is a heuristic and can change between days, garbling other accented names (è→č, ñ→ń, ò→ň …).Change
tap_uk_gias/gias_csv.py:read_gias_csv(content, logger)decodes the raw bytes as cp1252 and keeps every column a string, with blanks as''(as before).singer_sdkimport, so CI can test it.Verification
pipeline/tests/test_gias_encoding.py(4 tests): à, ’, é, ° and ç decode correctly, the windows-1250 reading isn't used, values stay strings, and an undefined byte is replaced and logged. All 4 failed before the change and pass after.GIASEstablishmentsStream.get_recordson the real 3 Oct extract, withresp.textforced to the windows-1250 reading: 52,586 rows, both St Thomas à Becket schools correct, no "ŕ", no replacement characters.python -m pytest backend/tests pipeline/tests scripts/ci/tests -q: 332 passed.After merge
The next daily GIAS run rewrites the names. The two schools above should then read "à" on the site.
Independent of #181 (scheduled dbt selectors). The IDACI and FBIT taps read the same extract the same way, apparently only for code columns, so they're left alone here.
🤖 Generated with Claude Code
🤖 AI Code Review (Claude Code)
The tap now decodes GIAS CSV bytes explicitly as Windows-1252 through a shared
read_gias_csvhelper, used by both the establishments and links streams. This replacesresp.text, where requests guessed the codec and turned "à" into "ŕ". Undefined bytes become U+FFFD and are logged instead of failing the load. The change is small and correct, and the tests cover the real byte sequences and the undefined-byte case. I found no blocking issues.🟡 Minor