GIAS publishes its CSVs in Windows-1252 and sends no charset. The tap read
resp.text, so requests guessed the codec, and the encoding="latin-1" passed
to read_csv did nothing on already-decoded text. On 3 Oct 2026 the guess was
windows-1250, and "St Thomas à Becket" (138950, 149557) was stored as
"St Thomas ŕ Becket". A different guess on another day would garble other
accented names.
Both streams now decode the downloaded bytes themselves (gias_csv.py). A byte
Windows-1252 leaves undefined becomes U+FFFD with a logged warning instead of
failing the load, so one odd name cannot stop the daily refresh. None of the
nine extracts checked (1 Jul to 3 Oct 2026) contains such a byte.
Checked by running the tap on the real 3 Oct extract with .text forced to
windows-1250: all 52,586 rows decode, with no "ŕ" and no replacement
characters.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The old sync created a collection, imported batches without reading a
single import response, and swapped the alias regardless. A partial
import published a half-empty index, and two overlapping DAG runs could
prune each other's collections.
Publication now checks every import response and the final document
count before upserting the alias, and holds a session-scoped advisory
lock across the read and the publish so concurrent runs serialise.
Cleanup keeps the previous collection as a rollback pointer and is
best-effort: an uncertain alias response must never delete what might
still be live.
Also parses the Typesense URL properly instead of splitting on colons,
which mangled any host carrying a scheme and a default port.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>