Compare commits

...
Author SHA1 Message Date
TudorandClaude Opus 5.5 94bfac9caf fix(pipeline): decode GIAS extracts as Windows-1252
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m16s
PR Checks / Backend Smoke (pull_request) Successful in 10s
PR Checks / Build Backend (no push) (pull_request) Successful in 18s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m27s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 1m17s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 19s
GIAS publishes its CSVs in Windows-1252 and sends no charset. The tap read
resp.text, so requests guessed the codec, and the encoding="latin-1" passed
to read_csv did nothing on already-decoded text. On 3 Oct 2026 the guess was
windows-1250, and "St Thomas à Becket" (138950, 149557) was stored as
"St Thomas ŕ Becket". A different guess on another day would garble other
accented names.

Both streams now decode the downloaded bytes themselves (gias_csv.py). A byte
Windows-1252 leaves undefined becomes U+FFFD with a logged warning instead of
failing the load, so one odd name cannot stop the daily refresh. None of the
nine extracts checked (1 Jul to 3 Oct 2026) contains such a byte.

Checked by running the tap on the real 3 Oct extract with .text forced to
windows-1250: all 52,586 rows decode, with no "ŕ" and no replacement
characters.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-04 10:09:39 +01:00
tudor 423b27140c Merge pull request 'fix(school): send the school page its admissions policy, and read it exactly' (#179) from fix/school-page-selective-flag into main
Stage (build -> staging -> E2E gate) / prepare (push) Successful in 1s
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 21s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 1m24s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 30s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Successful in 3m22s
Reviewed-on: #179
2026-10-02 22:56:36 +00:00
tudor 1c62e8247d Merge pull request 'fix(search): stop replaying a failed LA-averages request forever' (#178) from fix/la-average-cached-failure into main
Stage (build -> staging -> E2E gate) / prepare (push) Successful in 1s
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 19s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 1m27s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 27s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 3m47s
Reviewed-on: #178
2026-10-02 22:44:57 +00:00
TudorandClaude Opus 5.5 da5d63593f test(e2e): a no-faith comprehensive makes no Selective or faith claim
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m13s
PR Checks / Backend Smoke (pull_request) Successful in 10s
PR Checks / Build Backend (no push) (pull_request) Successful in 19s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m20s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 11s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 16s
The Admissions section of a non-selective secondary with no religious
character must say neither "Selective:" nor "Faith priority:". Run against
staging before the fix, it fails on "Faith priority" (Burntwood's "(None)").

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-02 23:43:43 +01:00
TudorandClaude Opus 5.5 ef4a2ccccb fix(school): send the school page its admissions policy, and read it exactly
The header's Selective flag read school_info.admissions_policy, which the
detail endpoint never sent, so no school page could flag Selective while
its search row did (staging E2E: The Grammar School at Leeds). The detail
payload now carries it, and a contract test checks it carries every field
the header's flags read.

Sending it would have switched on two older copies of the tag logic #176
fixed in the rows. The Admissions section and the cut-off note both tested
includes('selective'), so every non-selective secondary would have read
"entry is by selective examination". The section also counted "None" as a
faith: Burntwood reads "a faith-based admissions priority (None)" today.
All of them now share isSelective() and hasReligiousCharacter(), which also
treats "Not applicable" as no faith, as the place table already does.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-02 23:43:43 +01:00
TudorandClaude Opus 5.5 59ea8a4bdd fix(search): stop replaying a failed LA-averages request forever
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m12s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 18s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m18s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 11s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 18s
The search page fetched LA averages with cache: 'force-cache', which serves
any stored response, however old, without asking the server. One failed
request (a staging deploy restart; the July proxy outage) was stored and
replayed on every later visit, and the error was swallowed, so the
"vs LA avg" delta silently vanished from every secondary row in that
browser. A Playwright profile still held a 500 dated 5 July.

The default cache mode honours the API's Cache-Control (five minutes), so
a good answer is still reused and an error never is. Browsers holding a
stored failure recover on their next visit.

A journey now checks that a mainstream secondary's row shows the
comparison: nothing did, which is how it could go missing unnoticed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-02 23:20:14 +01:00
15 changed files with 318 additions and 34 deletions

No files matched your search

+2
View File
@@ -1047,6 +1047,8 @@ async def get_school_details(request: Request, urn: int):
"age_range": latest.get("age_range", ""),
"has_sixth_form": latest.get("has_sixth_form"),
"nursery_provision": latest.get("nursery_provision"),
# The header's Selective flag reads it (lib/schoolFacts).
"admissions_policy": latest.get("admissions_policy"),
"status": latest.get("status"),
"latitude": latest.get("latitude"),
"longitude": latest.get("longitude"),
@@ -0,0 +1,55 @@
"""The school page's header flags (nextjs-app/lib/schoolFacts.ts) read seven
school_info fields, so the detail payload must carry every one.
It lacked admissions_policy, so no school page could flag Selective while its
search row did: Tiffin and The Grammar School at Leeds showed the tag in
search and nothing on their own pages.
"""
import numpy as np
import pandas as pd
import pytest
from fastapi.testclient import TestClient
# The fields schoolFlags() picks from School (FlagFields in lib/schoolFacts.ts).
FLAG_FIELDS = (
"type_group", "admissions_policy", "gender", "religious_denomination",
"nursery_provision", "has_sixth_form", "phase",
)
def _school_df() -> pd.DataFrame:
return pd.DataFrame([{
"urn": 136910, "school_name": "Tiffin School", "phase": "Secondary",
"school_type": "Academy converter", "admissions_policy": "Selective",
"gender": "Boys", "religious_denomination": "Christian",
"nursery_provision": "Not applicable", "has_sixth_form": True,
"age_range": "11-18", "local_authority": "Kingston upon Thames",
"address": "Queen Elizabeth Road, Kingston upon Thames, KT2 6RL",
"postcode": "KT2 6RL", "latitude": 51.41, "longitude": -0.30,
"year": 202425, "attainment_8_score": 75.0, "total_pupils": 200,
"gias_total_pupils": 1478, "ofsted_grade": np.nan,
}])
@pytest.fixture()
def client(monkeypatch):
from backend import app as app_module
monkeypatch.setattr(app_module, "load_school_data", _school_df)
monkeypatch.setattr(app_module, "load_latest_school_data", _school_df)
monkeypatch.setattr(app_module, "get_supplementary_data", lambda db, urn: {})
monkeypatch.setattr(app_module, "_place_registry", None)
return TestClient(app_module.app, raise_server_exceptions=False)
def test_the_school_page_carries_every_field_its_flags_read(client):
resp = client.get("/api/schools/136910")
assert resp.status_code == 200, resp.text
info = resp.json()["school_info"]
assert [f for f in FLAG_FIELDS if f not in info] == []
def test_the_school_page_says_a_selective_school_is_selective(client):
info = client.get("/api/schools/136910").json()["school_info"]
assert info["admissions_policy"] == "Selective"
+51
View File
@@ -410,6 +410,29 @@ test('search and the school page agree on how many pupils a secondary has', asyn
expect(school.total_pupils).toBe(detail.school_info.total_pupils);
});
test('a secondary search row compares its Attainment 8 with the LA average', async ({ page }) => {
// The comparison vanished unnoticed: the averages were fetched with
// force-cache, so one stored failure hid it in that browser for good.
// Playwright disables the HTTP cache when it intercepts requests, so this
// guards the comparison itself; the unit test pins the cache mode.
const la = await (await page.request.get('/api/la-averages')).json();
const averages: Record<string, number> = la.secondary?.attainment_8_by_la ?? {};
const res = await page.request.get('/api/schools?search=school&phase=secondary&page_size=50');
expect(res.ok()).toBeTruthy();
const school = ((await res.json()).schools ?? []).find(
(s: { attainment_8_score?: number | null; local_authority?: string; school_type?: string }) =>
s.attainment_8_score != null && s.local_authority != null && averages[s.local_authority] != null
&& !/special|pupil referral|alternative provision/i.test(s.school_type ?? ''));
test.skip(!school, 'no mainstream secondary with an LA average here');
await searchByName(page, school.school_name);
const link = page.locator(`a[href^="/school/${school.urn}-"]`).first();
await expect(link).toBeVisible({ timeout: 15_000 });
const stats = link.locator('xpath=ancestor::div[contains(@class, "__rowContent")][1]')
.locator('[class*="__line3"]');
await expect(stats.getByText(/vs LA avg/)).toBeVisible();
});
test('a phase outside primary/secondary filters to that phase, not to everything', async ({ page }) => {
// The search page offers every GIAS phase, but the API only knew the grouped
// ones and silently dropped the rest — so "Nursery" returned primaries.
@@ -3396,3 +3419,31 @@ test('a selective school is flagged Selective, on its page and in search', async
const tags = await rowTags(page, school);
await expect(tags.getByText('Selective', { exact: true })).toBeVisible();
});
/*
* The Admissions section repeated the search rows' old tag logic: "selective"
* matched inside "Non-selective", and "None" counted as a faith, so Burntwood
* read "this school has a faith-based admissions priority (None)". The
* Selective half stayed hidden only because the school page's API did not
* send admissions_policy. Data-invariant: the school comes from the API, and
* must have an Admissions section to read.
*/
test('a non-selective school with no faith makes neither claim in its Admissions section', async ({ page }) => {
const res = await page.request.get(
'/api/schools?search=school&phase=secondary&faith=none&admissions_policy=non-selective&page_size=20');
expect(res.ok()).toBeTruthy();
let urn: number | null = null;
for (const s of ((await res.json()).schools ?? []).slice(0, 8)) {
if (s.religious_denomination !== 'None') continue;
const detail = await (await page.request.get(`/api/schools/${s.urn}`)).json();
if (detail.admissions) { urn = s.urn; break; }
}
test.skip(urn == null, 'no non-selective, no-faith secondary with admissions data here');
await page.goto(`/school/${urn}`);
const admissions = page.locator('section#admissions');
await expect(admissions).toBeVisible({ timeout: 15_000 });
await expect(admissions.getByText('Selective:', { exact: true })).toHaveCount(0);
await expect(admissions.getByText('Faith priority:', { exact: true })).toHaveCount(0);
await expect(admissions.getByText(/\(None\)/)).toHaveCount(0);
});
@@ -1,6 +1,6 @@
import { act, fireEvent, render, screen } from '@testing-library/react';
import { HomeView } from '@/components/HomeView';
import { fetchSchools } from '@/lib/api';
import { fetchLAaverages, fetchSchools } from '@/lib/api';
import { primaryFixture } from '../support/schoolFixtures';
import type { SchoolsResponse, School } from '@/lib/types';
@@ -84,3 +84,23 @@ test('failed map requests can be retried by reopening the map', async () => {
expect(fetchSchools).toHaveBeenCalledTimes(2);
expect(screen.getByTestId('map')).toHaveTextContent('Retry result');
});
test('LA averages are not fetched with force-cache, so one failure is not replayed for good', async () => {
// force-cache serves any stored response, however old, without asking the
// server. A request that failed once (a staging deploy restart, the July
// proxy outage) was stored and replayed on every later visit, and the
// "vs LA avg" delta vanished from every secondary row in that browser.
// The default mode honours the API's Cache-Control and never reuses an
// error.
params = new URLSearchParams('search=high');
const secondary: SchoolsResponse = {
...response('Alpha High'),
schools: [{ ...primaryFixture.schoolInfo, school_name: 'Alpha High', phase: 'Secondary', attainment_8_score: 50 }],
};
render(<HomeView initialSchools={secondary} filters={filters} />);
await act(async () => {});
expect(fetchLAaverages).toHaveBeenCalled();
for (const [options] of jest.mocked(fetchLAaverages).mock.calls) {
expect(options?.cache).not.toBe('force-cache');
}
});
@@ -0,0 +1,58 @@
/**
* The Admissions section's Selective and faith notes.
*
* It tested the policy with includes('selective'), which "Non-selective"
* passes, and excluded only "Does not apply" from the religious character, so
* Burntwood (no religious character, recorded as "None") read "this school
* has a faith-based admissions priority (None)". The Selective half never
* fired only because the school page's API did not send admissions_policy.
*/
import { screen, within } from '@testing-library/react';
import type { School } from '@/lib/types';
import { secondaryFixture } from '../support/schoolFixtures';
import { renderSecondarySchoolDetail } from '../support/renderSchoolDetail';
jest.mock('@/lib/analytics', () => ({
track: jest.fn(),
getNavigationSource: () => 'direct',
}));
jest.mock('@/components/PerformanceChart', () => ({
PerformanceChart: () => <div data-testid="performance-chart" />,
}));
jest.mock('@/components/AdmissionsTrendChart', () => ({
__esModule: true,
default: () => <div data-testid="admissions-trend-chart" />,
}));
jest.mock('@/components/SchoolHeroMap', () => ({
SchoolHeroMap: () => <div data-testid="hero-map" />,
__esModule: true,
}));
function admissionsOf(info: Partial<School>) {
renderSecondarySchoolDetail({ ...secondaryFixture, schoolInfo: { ...secondaryFixture.schoolInfo, ...info } });
return within(screen.getByRole('heading', { name: 'Admissions' }).closest('section')!);
}
describe('Admissions section notes', () => {
it('notes the entrance test for a selective school', () => {
const section = admissionsOf({ admissions_policy: 'Selective', religious_denomination: 'None' });
expect(section.getByText('Selective:')).toBeInTheDocument();
});
it('does not call a non-selective school selective', () => {
const section = admissionsOf({ admissions_policy: 'Non-selective', religious_denomination: 'None' });
expect(section.queryByText('Selective:')).not.toBeInTheDocument();
});
it.each(['None', 'Does not apply'])('claims no faith priority when the register says %p', (religious_denomination) => {
// Not "Non-selective": a misread Selective would win and hide this case.
const section = admissionsOf({ admissions_policy: 'Not applicable', religious_denomination });
expect(section.queryByText('Faith priority:')).not.toBeInTheDocument();
});
it('notes the religious character of a faith school', () => {
const section = admissionsOf({ admissions_policy: 'Non-selective', religious_denomination: 'Church of England' });
expect(section.getByText('Faith priority:')).toBeInTheDocument();
});
});
@@ -157,6 +157,13 @@ describe('compareToCutoff', () => {
});
describe('describeCutoffAbsence', () => {
it('does not read "Non-selective" as selective', () => {
// A substring test matched "selective" inside "Non-selective".
const s = describeCutoffAbsence({ localAuthority: 'Wandsworth', admissionsPolicy: 'Non-selective' });
expect(s).not.toMatch(/entrance test/);
expect(s).toMatch(/^Wandsworth has not published/);
});
it('explains a selective school by how it admits, not as missing data', () => {
const s = describeCutoffAbsence({ localAuthority: 'Kent', admissionsPolicy: 'Selective' });
expect(s).toContain('entrance test');
@@ -59,6 +59,8 @@ describe('schoolFlags', () => {
it('never prints Non-selective, and never a no-faith value', () => {
expect(labels({ admissions_policy: 'Non-selective', religious_denomination: 'Does not apply' })).toEqual([]);
expect(labels({ admissions_policy: 'Not applicable', religious_denomination: 'None' })).toEqual([]);
// PlaceView's rule: GIAS sometimes says "Not applicable" for no faith too.
expect(labels({ religious_denomination: 'Not applicable' })).toEqual([]);
});
it('prints the religious character as the register records it', () => {
+6 -2
View File
@@ -358,10 +358,14 @@ export function HomeView({ initialSchools, filters, totalSchools, howItWorks, ed
return () => controller.abort();
}, [resultsView, searchParams, initialSchools.schools]);
// Fetch LA averages when secondary or mixed schools are visible
// Fetch LA averages when secondary or mixed schools are visible. Default
// cache mode, never force-cache: force-cache replays any stored response
// without asking the server, so one failed request (a deploy restart) hid
// every "vs LA avg" delta in that browser for good. The API's Cache-Control
// already lets the browser reuse a good answer for five minutes.
useEffect(() => {
if (!isSecondaryView && !isMixedView) return;
fetchLAaverages({ cache: 'force-cache' })
fetchLAaverages()
.then(data => setLaAverages(data.secondary.attainment_8_by_la))
.catch(() => {});
}, [isSecondaryView, isMixedView]);
@@ -7,7 +7,7 @@
*/
import type { School, SchoolAdmissions, SchoolAdmissionDistance } from '@/lib/types';
import { formatPercentage } from '@/lib/utils';
import { formatPercentage, hasReligiousCharacter, isSelective } from '@/lib/utils';
import { Section, sectionStyles as styles } from './sectionShared';
import {
describeCutoff, describeCutoffAbsence,
@@ -33,10 +33,8 @@ export function SecondaryAdmissionsSection({
const featureOn = admissionDistance !== undefined;
// Moved with this section from SecondarySchoolDetailView, its only consumer.
const admissionsTag = (() => {
const policy = schoolInfo.admissions_policy?.toLowerCase() ?? '';
if (policy.includes('selective')) return 'Selective';
const denom = schoolInfo.religious_denomination ?? '';
if (denom && denom !== 'Does not apply') return 'Faith priority';
if (isSelective(schoolInfo.admissions_policy)) return 'Selective';
if (hasReligiousCharacter(schoolInfo.religious_denomination)) return 'Faith priority';
return null;
})();
@@ -16,7 +16,7 @@
*/
import type { SchoolAdmissionDistance } from '@/lib/types';
import { formatCutoffDistance, formatMiles, formatEntryYear } from '@/lib/utils';
import { formatCutoffDistance, formatMiles, formatEntryYear, isSelective } from '@/lib/utils';
export interface CutoffDisplay {
/** Headline figure, e.g. "0.31 miles". */
@@ -193,8 +193,7 @@ export function describeCutoffAbsence({
admissionsPolicy,
admissionsHistory = [],
}: AbsenceInput): string {
const policy = (admissionsPolicy ?? '').toLowerCase();
if (policy.includes('selective')) {
if (isSelective(admissionsPolicy)) {
return 'Places at this school are ranked by the entrance test rather than by '
+ 'distance, so no cut-off distance applies.';
}
+2 -2
View File
@@ -7,7 +7,7 @@
*/
import type { School } from './types';
import { hasNurseryClasses, hasReligiousCharacter, singleSexLabel } from './utils';
import { hasNurseryClasses, hasReligiousCharacter, isSelective, singleSexLabel } from './utils';
/** The search filter's type groups (backend/school_groups.py), without its
* parenthesised notes: fees have a flag of their own. */
@@ -54,7 +54,7 @@ export function schoolFlags(school: FlagFields): SchoolFlag[] {
const phase = school.phase?.trim().toLowerCase();
if (school.type_group === 'independent') condition('Fee-paying');
if (school.admissions_policy?.trim().toLowerCase() === 'selective') condition('Selective');
if (isSelective(school.admissions_policy)) condition('Selective');
const singleSex = singleSexLabel(school.gender);
if (singleSex) condition(singleSex);
if (hasReligiousCharacter(school.religious_denomination)) {
+13 -3
View File
@@ -923,12 +923,22 @@ export function isSpecialSchool(school: { school_type?: string | null }): boolea
}
/**
* Whether GIAS records a religious character. "None" and "Does not apply" are
* the register's two ways of saying it has none, and neither is a faith.
* Whether GIAS records a religious character. "None", "Does not apply" and
* "Not applicable" are the register's ways of saying it has none; the place
* table (PlaceView's NO_FAITH) reads the same three.
*/
export function hasReligiousCharacter(value: string | null | undefined): boolean {
const v = value?.trim().toLowerCase() ?? '';
return v !== '' && v !== 'none' && v !== 'does not apply';
return v !== '' && v !== 'none' && v !== 'does not apply' && v !== 'not applicable';
}
/**
* Whether GIAS records the school as selective. Exact: "Non-selective"
* contains "selective", so a substring test read every comprehensive as
* selective. (The register files a partly selective school as non-selective.)
*/
export function isSelective(admissionsPolicy: string | null | undefined): boolean {
return admissionsPolicy?.trim().toLowerCase() === 'selective';
}
/**
@@ -0,0 +1,26 @@
"""Read a GIAS extract from the raw bytes of the download.
GIAS writes its CSVs in Windows-1252 and sends no charset, so `resp.text`
leaves requests to guess the codec. On 3 Oct 2026 it guessed windows-1250 and
"à" became "ŕ". Decode the bytes ourselves instead.
"""
from __future__ import annotations
import io
import pandas as pd
GIAS_ENCODING = "cp1252"
def read_gias_csv(content: bytes, logger=None) -> pd.DataFrame:
"""Every column as a string; a blank cell stays ''."""
# Windows-1252 leaves five bytes undefined. One stray byte must not stop
# the daily refresh of every school, so it becomes U+FFFD and is logged.
text = content.decode(GIAS_ENCODING, errors="replace")
undecodable = text.count("�")
if undecodable and logger is not None:
logger.warning("%d byte(s) in the GIAS extract could not be decoded as %s",
undecodable, GIAS_ENCODING)
return pd.read_csv(io.StringIO(text), dtype=str, keep_default_na=False)
@@ -7,6 +7,8 @@ from datetime import date, timedelta
from singer_sdk import Stream, Tap
from singer_sdk import typing as th
from tap_uk_gias.gias_csv import read_gias_csv
GIAS_URL_TEMPLATE = (
"https://ea-edubase-api-prod.azurewebsites.net"
"/edubase/downloads/public/edubasealldata{date}.csv"
@@ -74,9 +76,6 @@ class GIASEstablishmentsStream(Stream):
def get_records(self, context):
"""Download GIAS CSV and yield rows."""
import io
import pandas as pd
import requests
today = date.today()
@@ -94,12 +93,7 @@ class GIASEstablishmentsStream(Stream):
resp.raise_for_status()
df = pd.read_csv(
io.StringIO(resp.text),
encoding="latin-1",
dtype=str,
keep_default_na=False,
)
df = read_gias_csv(resp.content, self.logger)
for _, row in df.iterrows():
record = row.to_dict()
@@ -126,9 +120,6 @@ class GIASLinksStream(Stream):
def get_records(self, context):
"""Download GIAS links CSV and yield rows."""
import io
import pandas as pd
import requests
today = date.today()
@@ -146,12 +137,7 @@ class GIASLinksStream(Stream):
resp.raise_for_status()
df = pd.read_csv(
io.StringIO(resp.text),
encoding="latin-1",
dtype=str,
keep_default_na=False,
)
df = read_gias_csv(resp.content, self.logger)
for _, row in df.iterrows():
record = row.to_dict()
+66
View File
@@ -0,0 +1,66 @@
"""GIAS publishes its extracts in Windows-1252 and declares no charset.
The tap used to hand pandas `resp.text`, so requests guessed the codec.
On 3 Oct 2026 it guessed windows-1250, and "St Thomas à Becket" was stored
as "St Thomas ŕ Becket". The `encoding=` passed to read_csv did nothing,
because the text was already decoded.
"""
import importlib.util
import logging
from pathlib import Path
import pytest
MODULE = (Path(__file__).resolve().parents[1] / 'plugins' / 'extractors' / 'tap-uk-gias'
/ 'tap_uk_gias' / 'gias_csv.py')
@pytest.fixture
def gias_csv():
spec = importlib.util.spec_from_file_location('gias_csv', MODULE)
module = importlib.util.module_from_spec(spec)
spec.loader.exec_module(module)
return module
# Byte for byte as GIAS writes it: 0xE0 à, 0x92 ’, 0xE9 é, 0xB0 °, 0xE7 ç.
EXTRACT = (
b'"URN","EstablishmentName","HeadLastName"\r\n'
b'"138950","St Thomas \xe0 Becket Catholic Secondary School","Smith"\r\n'
b'"100000","The Dean and Chapter of St Paul\x92s Cathedral","Pr\xe9vert"\r\n'
b'"140677","North Star 180\xb0","Fran\xe7ois"\r\n'
b'"100001","No head recorded",""\r\n'
)
def test_names_decode_as_windows_1252(gias_csv):
df = gias_csv.read_gias_csv(EXTRACT)
assert list(df['EstablishmentName']) == [
'St Thomas à Becket Catholic Secondary School',
'The Dean and Chapter of St Paul’s Cathedral',
'North Star 180°',
'No head recorded',
]
assert list(df['HeadLastName']) == ['Smith', 'Prévert', 'François', '']
def test_the_codec_requests_guessed_is_not_used(gias_csv):
# What the tap stored on 3 Oct: the same bytes read as windows-1250.
assert 'ŕ' in EXTRACT.decode('cp1250')
names = ' '.join(gias_csv.read_gias_csv(EXTRACT)['EstablishmentName'])
assert 'ŕ' not in names
def test_values_stay_strings(gias_csv):
df = gias_csv.read_gias_csv(EXTRACT)
assert df.loc[0, 'URN'] == '138950'
def test_a_byte_windows_1252_leaves_undefined_does_not_stop_the_load(gias_csv, caplog):
# 0x81 has no Windows-1252 character. One odd name must not block the daily
# refresh of every school, but it must be visible in the log.
extract = b'"URN","EstablishmentName"\r\n"100002","Odd \x81 Name"\r\n'
with caplog.at_level(logging.WARNING):
df = gias_csv.read_gias_csv(extract, logger=logging.getLogger('gias'))
assert df.loc[0, 'EstablishmentName'] == 'Odd � Name'
assert 'could not be decoded' in caplog.text