The em dash is one of the clearest tells of machine-written text, which is the exact impression this work exists to remove. Rewritten rather than substituted: where a dash was carrying a real aside the sentence is split or recast, not patched with a comma. Covers the About page, the two Callout labels an editor sees in the admin panel, and PUBLISHING.md, which defines the house style and should follow it. The rule is now recorded in that house style and in the spec's voice rules, so it survives this branch. Code comments are left alone: they are not copy, and the surrounding codebase uses the same punctuation throughout. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017YmbBhr8s7GusjDE12hrZM
17 KiB
Giving schoolcompare a human author: an About page and a blog
Date: 2026-09-02
Status: Design — awaiting review
Scope: A named author for the site, an /about page, and a Payload-CMS-backed
blog at /blog.
Why
The site reads as synthetic. Not because of its tone, but because of three specific absences:
- Nobody is accountable for the numbers. There is no author, no statement
of why the site exists, and no one who can be wrong. The only human trace on
the entire site is
contact@schoolcompare.co.ukin the footer. - No visible judgement. Every figure is presented as though it fell out of
a machine. Hundreds of editorial decisions went into this codebase — which
metrics to show, when a benchmark is invalid, what to suppress — and not one
of them is visible to a reader.
isSpecialSchool()silently drops the England comparison for special schools and PRUs because that comparison is meaningless; nowhere does the site say so. - The voice is institutional third person. "schoolcompare brings it all into one place." "Built for parents, governors, journalists." That is brochure register, and it is precisely the register that machine-generated content defaults to.
There is a second, independent reason. The SEO programme
(2026-08-20-seo-programme-design.md) defines eight workstreams and none of
them address E-E-A-T or authorship. School performance data is YMYL territory;
an anonymous site republishing DfE figures has no authorship signal at all. This
work fills that hole, and the blog gives W6 (explainer content) somewhere to
live.
The failure mode to avoid
The standard fix — a stock photo and "Hi, I'm Tudor, and I'm passionate about education!" — reads as more synthetic than the current coldness. Manufactured warmth is a stronger machine-tell than plain institutional voice. Everything here has to be specific, occasionally awkward, and willing to be unflattering, or it makes the problem worse.
Positioning
The author is Tudor: first name only, real photograph, no surname, no employer named.
The credibility claim is deliberately not educational expertise. The About page states plainly: "I'm not an education expert." Authority comes from two things that are actually true:
- Experience. A parent going through primary admissions in south-west London right now. Google's E-E-A-T leads with Experience, and lived experience of the thing is exactly what the DfE's own service lacks.
- Method. Every number's provenance is stated, so a reader can check the site rather than trust it.
This is more durable than borrowed expertise: it cannot be undermined by someone noticing the author has no teaching qualification.
Consequence for the design. A Person entity with no surname is a weak
search signal and cannot be corroborated off-site. The credibility load
therefore shifts onto the methodology being visibly rigorous. That is a design
constraint, not a caveat — it is why the About page carries a substantial
"how this is built and where it can be wrong" section rather than a short bio.
Voice rules
Applied to About and every post. Recorded here so the voice does not drift.
- First person singular. "I built", not "we provide".
- Concrete over general. "when we were looking at schools in Wandsworth" beats any amount of stated warmth.
- State limits before someone else finds them. Every post that presents a metric says what it does not show.
- No mission statements, no "passionate about", no invented team.
- No em dashes. One of the clearest tells of machine-written prose, which is the exact problem this work exists to fix.
- Short sentences. The existing code comments in this repo are already written this way; the prose should match.
Scope
In:
/about— a coded page (not CMS-managed)./blogand/blog/[slug]— Payload-backed, with an index and post pages.- Payload CMS installed into the existing Next application.
- Footer and navigation links to both.
Person,Organization,BlogPosting,BreadcrumbListJSON-LD.- RSS feed and sitemap integration.
- One first post, so the blog does not launch empty.
Out (deliberately):
- Rewriting existing homepage/how-it-works copy into first person. Worth doing, but it would double the review surface of this PR. Separate change.
- In-product signed notes on school pages (the "distributed humanity" idea). Revisit once About and the blog exist.
- Comments, newsletter, author accounts beyond one.
- A team page. There is no team.
Architecture
Topology
Payload 3 installs into the existing Next application and serves /admin
from the same container. One image, one deploy, no new service. This is
Payload 3's native model and it makes on-demand revalidation trivial, because
the CMS hooks run in the same process as the Next cache.
Accepted costs: the public site's image now carries Payload, so a CMS security patch redeploys the whole site; and the image grows substantially.
Two collisions that must be handled
1. /api is already taken. app/api/[...path]/route.ts is a catch-all that
proxies /api/* to FastAPI at runtime. Payload's default API route is also
/api. Left alone, these fight, and the failure is not clean — the catch-all
would swallow Payload's admin API calls and forward them to FastAPI.
Payload's API route is therefore remapped:
routes: { api: '/cms-api', admin: '/admin' }
with its route group at app/(payload)/cms-api/[...slug]/route.ts. The
/cms-api prefix must also be added to the FastAPI proxy's excluded-paths list
as a defensive second line.
2. next.config.js is CommonJS. Payload's withPayload() wrapper is ESM
only. The config must become next.config.mjs, converting module.exports to
export default and wrapping the export. All existing content — the standalone
output, outputFileTracingIncludes, the staging X-Robots-Tag header block,
the CSP — carries over unchanged. This is mechanical but it touches the file
that controls staging's noindex, so it needs care and an explicit test.
Database
Payload uses the existing sc_database Postgres instance, in its own
payload schema:
db: postgresAdapter({
pool: { connectionString: process.env.DATABASE_URL },
schemaName: 'payload',
})
The frontend container is already on the backend Docker network, so it can
reach sc_database:5432 with no networking change. It needs a new
DATABASE_URL environment variable.
Schema isolation is not cosmetic. public currently holds the application
tables and Airflow's metadata, and scripts/migrate_csv_to_db.py --drop exists
to drop and reimport. Blog content living in its own schema means no data
pipeline operation can destroy it.
Verified 2026-09-02 (this was an open question when the spec was written).
--drop calls run_full_migration() in backend/migration.py, which drops
exactly two tables by name:
ks2_tables = ["school_results", "schools"]
for tname in ks2_tables:
if tname in existing:
Base.metadata.tables[tname].drop(bind=engine)
There is no Base.metadata.drop_all() anywhere in backend/, and no
DROP SCHEMA. The only other drop is _apply_schema_drops(), a single
schema-qualified DROP TABLE IF EXISTS marts.fact_parent_view CASCADE.
Nothing sets search_path, so the SQLAlchemy metadata resolves to public,
and inspector.get_table_names() does not even enumerate other schemas.
So the guarantee is stronger than schema isolation alone: --drop targets two
named tables that Payload does not have, and would not reach posts, media
or users even if they shared a schema. The payload schema remains the right
choice — it protects against a future broadening of that script rather than
today's behaviour — but the safety claim rests on verified code, not on
assumption.
Putting CMS tables in this instance is consistent with existing practice — Airflow already stores its metadata there.
Migrations
Payload's Postgres adapter auto-pushes schema in development and requires
explicit migrations in production. Use prodMigrations, which runs pending
migrations during server initialisation:
db: postgresAdapter({ /* ... */, prodMigrations: migrations })
This is preferred over a one-shot init container (the airflow-init pattern)
because the app is a single long-running process and there is no ordering
problem to solve. Migration files are generated with payload migrate:create
and committed, so schema changes travel through the same PR and staging gate as
code.
Media
Uploads go to a Docker named volume, consistent with postgres_data,
typesense_data and airflow_logs.
staticDirmust be an absolute path in Payload 3:/app/media.- The container runs as
nextjs(uid 1001). The Dockerfile mustmkdir -p /app/media && chown nextjs:nodejs /app/mediabefore the volume is mounted, or Docker will create the mountpoint root-owned and every upload will fail with EACCES. sharpmoves fromdevDependenciestodependencies— Payload needs it at runtime to generateimageSizes.- The volume must be added to the backup routine alongside Postgres. A blog post's images are not reproducible from the pipeline.
Rendering
Constraint: CI builds the image with no database reachable. Blog pages
therefore cannot use build-time generateStaticParams — that would either fail
the build or bake in an empty post list.
Instead: ISR. Post and index pages declare a revalidate window and render on
first request, with Payload afterChange / afterDelete hooks calling
revalidatePath('/blog') and revalidatePath('/blog/' + slug) for immediate
publication. Because Payload runs in the same process, the hook calls
revalidatePath from next/cache directly — no webhook, no shared secret.
The ISR cache lives on container disk and is cleared by a redeploy. For a single container serving a handful of posts this is fine.
Collections
posts—title,slug,publishedAt,excerpt,heroImage(relation tomedia),content(Lexical rich text),seogroup (metaTitle,metaDescription),_status(drafts enabled).media— upload collection,altrequired,imageSizesfor thumbnail and hero widths, public read access.users— Payload's auth collection. One account. Public creation disabled.
Drafts are enabled so posts can be written over several sittings and previewed before publication.
Payload Blocks are how posts embed live product components — a real trend chart or comparison table inside a post, rendered from live data rather than screenshotted. This is the main thing the CMS has to earn back against file-based MDX, and it directly serves the goal: showing judgement in context. Ship with one block (a callout/aside for "what this number doesn't tell you"); add a live-chart block once a post needs it.
Security
/admin is the first authenticated surface on this site. Public, hardened:
PAYLOAD_SECRET— long, random, set in the Portainer stack environment, never committed. The same variable must exist in staging with a different value.- Strong unique password on the single admin account.
- Login rate limiting via Payload's
maxLoginAttempts/lockTime. X-Robots-Tag: noindex, nofollowon/admin/*and/cms-api/*, and arobots.tsdisallow. The admin panel must never be indexed.- Public user creation disabled; no open registration.
- Verify the existing CSP
frame-ancestorsdirective does not break the admin panel.
Residual risk, accepted: a future Payload authentication CVE is live against the
public internet. Mitigation is prompt patching, which the staging→prod pipeline
already supports. If this becomes uncomfortable, restricting /admin at the
proxy to LAN/VPN is a one-line change later.
Staging note: staging runs the same image on stx., so it gets its own admin
panel and its own database. It must have its own PAYLOAD_SECRET and its own
credentials — never production's.
Deployment changes
nextjs-app/Dockerfile— create and chown/app/media; ensure Payload's admin bundle andsharpsurvive standalone output file tracing.docker-compose.portainer.ymland the staging equivalent — addDATABASE_URLandPAYLOAD_SECRETto thefrontendservice, add apayload_mediavolume mounted at/app/media, and adddepends_on: sc_database.- Document both new environment variables in the compose header comment block, which is where this stack records its configuration.
SEO
Person(Tudor, with photo) andOrganizationJSON-LD on/about.BlogPosting+BreadcrumbListon post pages, withauthorreferencing the samePerson.- Canonical URLs on
/blogand every post. - Posts and
/aboutadded to the existing sitemap (app/sitemap.xml/route.tsandapp/sitemaps/[...parts]). Post URLs come from Payload at request time. - RSS feed at
/blog/rss.xml. - Footer links to both pages, under a new "About" column.
Navigation is deliberately left alone. Navigation.tsx renders a bottom tab
bar on mobile that already carries four items (Search, Compare, Rankings,
Admissions). A fifth tab makes each one cramped at 320px, and About and Blog are
both lower-intent than any of the four. Both live in the footer; About
additionally gets a byline link from every post, which is where a reader who
cares actually asks the question. Revisit only if analytics show people hunting
for it.
Testing
Unit (Jest):
- Post rendering, including a post with no hero image and one with no excerpt.
- Slug generation and collision handling.
- JSON-LD shape for
BlogPostingandPerson. - The
next.config.mjsconversion preserves the stagingX-Robots-Tagrule — this guards the riskiest mechanical change in the plan.
E2E (Playwright, e2e/, required by CLAUDE.md for user-facing change):
/aboutrenders, shows the author name and photo, and is reachable from the footer and nav./bloglists at least one post; clicking through reaches the post.- A post page renders title, date, body and byline.
/adminresponds withnoindexand does not leak a stack trace when unauthenticated.
Note the known constraint: new journeys cannot be proven in PR checks, because the staging E2E gate runs post-merge.
Risks
| Risk | Mitigation |
|---|---|
next.config.mjs conversion silently drops the staging noindex header, making staging a crawlable duplicate |
Unit test asserting the header rule; verify on staging before promotion |
Payload API route collides with the FastAPI /api proxy |
Remap to /cms-api; add to the proxy's exclusion list |
| Media volume mounts root-owned; all uploads fail with EACCES | mkdir+chown in the Dockerfile before the mount; test an upload on staging |
| Build fails or bakes empty content because CI has no DB | No build-time DB access; ISR only |
A pipeline --drop destroys blog content |
Separate payload schema; verify --drop blast radius before building |
| Media volume not backed up; images unrecoverable | Add payload_media to the backup routine |
| Payload auth CVE exposed publicly | Prompt patching; proxy restriction available as a fallback |
| Blog launches empty or goes stale | Ship with one post; cadence is explicitly "a few times a year", so no cadence is promised anywhere on the page — no dates implying a schedule |
Sequence
Each step is independently reviewable and mergeable.
- Payload foundation — install,
next.config.mjsconversion,payloadschema,/cms-apiremap,userscollection,/adminhardening, compose and Dockerfile changes. No public-facing change yet. Verify on staging that the site is unchanged and/adminworks. /about— coded page, photo,Person/OrganizationJSON-LD, footer and nav links, e2e journey. Independently valuable and does not depend on the blog.- Blog —
postsandmediacollections,/blogindex and post pages, ISR plus revalidation hooks, RSS, sitemap, structured data, e2e journeys. - First post — written in the admin panel, published through the normal flow, proving the whole path end to end.
Step 1 carries all the infrastructure risk and none of the visible benefit, so it should be verified on staging carefully before step 2 starts.
Dependencies on Tudor
- A photograph. Blocks step 2. Nothing else in the plan is blocked by it.
- The first post's subject. Blocks step 4 only. Suggested: what school performance data cannot tell you — it demonstrates judgement, is genuinely useful, and is the kind of thing an anonymous or machine-written site will not publish.
Confirmation thatResolved 2026-09-02 — verified inscripts/migrate_csv_to_db.py --dropis schema-scoped.backend/migration.py; see the Database section. No action needed.