Files
school_compare/docs/superpowers/specs/2026-09-02-about-and-blog-design.md
T
TudorandClaude Opus 5 3f3c5953f6 style(copy): remove em dashes from the site's prose
The em dash is one of the clearest tells of machine-written text, which
is the exact impression this work exists to remove. Rewritten rather
than substituted: where a dash was carrying a real aside the sentence is
split or recast, not patched with a comma.

Covers the About page, the two Callout labels an editor sees in the
admin panel, and PUBLISHING.md, which defines the house style and should
follow it. The rule is now recorded in that house style and in the
spec's voice rules, so it survives this branch.

Code comments are left alone: they are not copy, and the surrounding
codebase uses the same punctuation throughout.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017YmbBhr8s7GusjDE12hrZM
2026-09-02 18:33:27 +01:00

17 KiB

Giving schoolcompare a human author: an About page and a blog

Date: 2026-09-02 Status: Design — awaiting review Scope: A named author for the site, an /about page, and a Payload-CMS-backed blog at /blog.

Why

The site reads as synthetic. Not because of its tone, but because of three specific absences:

  1. Nobody is accountable for the numbers. There is no author, no statement of why the site exists, and no one who can be wrong. The only human trace on the entire site is contact@schoolcompare.co.uk in the footer.
  2. No visible judgement. Every figure is presented as though it fell out of a machine. Hundreds of editorial decisions went into this codebase — which metrics to show, when a benchmark is invalid, what to suppress — and not one of them is visible to a reader. isSpecialSchool() silently drops the England comparison for special schools and PRUs because that comparison is meaningless; nowhere does the site say so.
  3. The voice is institutional third person. "schoolcompare brings it all into one place." "Built for parents, governors, journalists." That is brochure register, and it is precisely the register that machine-generated content defaults to.

There is a second, independent reason. The SEO programme (2026-08-20-seo-programme-design.md) defines eight workstreams and none of them address E-E-A-T or authorship. School performance data is YMYL territory; an anonymous site republishing DfE figures has no authorship signal at all. This work fills that hole, and the blog gives W6 (explainer content) somewhere to live.

The failure mode to avoid

The standard fix — a stock photo and "Hi, I'm Tudor, and I'm passionate about education!" — reads as more synthetic than the current coldness. Manufactured warmth is a stronger machine-tell than plain institutional voice. Everything here has to be specific, occasionally awkward, and willing to be unflattering, or it makes the problem worse.

Positioning

The author is Tudor: first name only, real photograph, no surname, no employer named.

The credibility claim is deliberately not educational expertise. The About page states plainly: "I'm not an education expert." Authority comes from two things that are actually true:

  • Experience. A parent going through primary admissions in south-west London right now. Google's E-E-A-T leads with Experience, and lived experience of the thing is exactly what the DfE's own service lacks.
  • Method. Every number's provenance is stated, so a reader can check the site rather than trust it.

This is more durable than borrowed expertise: it cannot be undermined by someone noticing the author has no teaching qualification.

Consequence for the design. A Person entity with no surname is a weak search signal and cannot be corroborated off-site. The credibility load therefore shifts onto the methodology being visibly rigorous. That is a design constraint, not a caveat — it is why the About page carries a substantial "how this is built and where it can be wrong" section rather than a short bio.

Voice rules

Applied to About and every post. Recorded here so the voice does not drift.

  • First person singular. "I built", not "we provide".
  • Concrete over general. "when we were looking at schools in Wandsworth" beats any amount of stated warmth.
  • State limits before someone else finds them. Every post that presents a metric says what it does not show.
  • No mission statements, no "passionate about", no invented team.
  • No em dashes. One of the clearest tells of machine-written prose, which is the exact problem this work exists to fix.
  • Short sentences. The existing code comments in this repo are already written this way; the prose should match.

Scope

In:

  • /about — a coded page (not CMS-managed).
  • /blog and /blog/[slug] — Payload-backed, with an index and post pages.
  • Payload CMS installed into the existing Next application.
  • Footer and navigation links to both.
  • Person, Organization, BlogPosting, BreadcrumbList JSON-LD.
  • RSS feed and sitemap integration.
  • One first post, so the blog does not launch empty.

Out (deliberately):

  • Rewriting existing homepage/how-it-works copy into first person. Worth doing, but it would double the review surface of this PR. Separate change.
  • In-product signed notes on school pages (the "distributed humanity" idea). Revisit once About and the blog exist.
  • Comments, newsletter, author accounts beyond one.
  • A team page. There is no team.

Architecture

Topology

Payload 3 installs into the existing Next application and serves /admin from the same container. One image, one deploy, no new service. This is Payload 3's native model and it makes on-demand revalidation trivial, because the CMS hooks run in the same process as the Next cache.

Accepted costs: the public site's image now carries Payload, so a CMS security patch redeploys the whole site; and the image grows substantially.

Two collisions that must be handled

1. /api is already taken. app/api/[...path]/route.ts is a catch-all that proxies /api/* to FastAPI at runtime. Payload's default API route is also /api. Left alone, these fight, and the failure is not clean — the catch-all would swallow Payload's admin API calls and forward them to FastAPI.

Payload's API route is therefore remapped:

routes: { api: '/cms-api', admin: '/admin' }

with its route group at app/(payload)/cms-api/[...slug]/route.ts. The /cms-api prefix must also be added to the FastAPI proxy's excluded-paths list as a defensive second line.

2. next.config.js is CommonJS. Payload's withPayload() wrapper is ESM only. The config must become next.config.mjs, converting module.exports to export default and wrapping the export. All existing content — the standalone output, outputFileTracingIncludes, the staging X-Robots-Tag header block, the CSP — carries over unchanged. This is mechanical but it touches the file that controls staging's noindex, so it needs care and an explicit test.

Database

Payload uses the existing sc_database Postgres instance, in its own payload schema:

db: postgresAdapter({
  pool: { connectionString: process.env.DATABASE_URL },
  schemaName: 'payload',
})

The frontend container is already on the backend Docker network, so it can reach sc_database:5432 with no networking change. It needs a new DATABASE_URL environment variable.

Schema isolation is not cosmetic. public currently holds the application tables and Airflow's metadata, and scripts/migrate_csv_to_db.py --drop exists to drop and reimport. Blog content living in its own schema means no data pipeline operation can destroy it.

Verified 2026-09-02 (this was an open question when the spec was written). --drop calls run_full_migration() in backend/migration.py, which drops exactly two tables by name:

ks2_tables = ["school_results", "schools"]
for tname in ks2_tables:
    if tname in existing:
        Base.metadata.tables[tname].drop(bind=engine)

There is no Base.metadata.drop_all() anywhere in backend/, and no DROP SCHEMA. The only other drop is _apply_schema_drops(), a single schema-qualified DROP TABLE IF EXISTS marts.fact_parent_view CASCADE. Nothing sets search_path, so the SQLAlchemy metadata resolves to public, and inspector.get_table_names() does not even enumerate other schemas.

So the guarantee is stronger than schema isolation alone: --drop targets two named tables that Payload does not have, and would not reach posts, media or users even if they shared a schema. The payload schema remains the right choice — it protects against a future broadening of that script rather than today's behaviour — but the safety claim rests on verified code, not on assumption.

Putting CMS tables in this instance is consistent with existing practice — Airflow already stores its metadata there.

Migrations

Payload's Postgres adapter auto-pushes schema in development and requires explicit migrations in production. Use prodMigrations, which runs pending migrations during server initialisation:

db: postgresAdapter({ /* ... */, prodMigrations: migrations })

This is preferred over a one-shot init container (the airflow-init pattern) because the app is a single long-running process and there is no ordering problem to solve. Migration files are generated with payload migrate:create and committed, so schema changes travel through the same PR and staging gate as code.

Media

Uploads go to a Docker named volume, consistent with postgres_data, typesense_data and airflow_logs.

  • staticDir must be an absolute path in Payload 3: /app/media.
  • The container runs as nextjs (uid 1001). The Dockerfile must mkdir -p /app/media && chown nextjs:nodejs /app/media before the volume is mounted, or Docker will create the mountpoint root-owned and every upload will fail with EACCES.
  • sharp moves from devDependencies to dependencies — Payload needs it at runtime to generate imageSizes.
  • The volume must be added to the backup routine alongside Postgres. A blog post's images are not reproducible from the pipeline.

Rendering

Constraint: CI builds the image with no database reachable. Blog pages therefore cannot use build-time generateStaticParams — that would either fail the build or bake in an empty post list.

Instead: ISR. Post and index pages declare a revalidate window and render on first request, with Payload afterChange / afterDelete hooks calling revalidatePath('/blog') and revalidatePath('/blog/' + slug) for immediate publication. Because Payload runs in the same process, the hook calls revalidatePath from next/cache directly — no webhook, no shared secret.

The ISR cache lives on container disk and is cleared by a redeploy. For a single container serving a handful of posts this is fine.

Collections

  • posts — title, slug, publishedAt, excerpt, heroImage (relation to media), content (Lexical rich text), seo group (metaTitle, metaDescription), _status (drafts enabled).
  • media — upload collection, alt required, imageSizes for thumbnail and hero widths, public read access.
  • users — Payload's auth collection. One account. Public creation disabled.

Drafts are enabled so posts can be written over several sittings and previewed before publication.

Payload Blocks are how posts embed live product components — a real trend chart or comparison table inside a post, rendered from live data rather than screenshotted. This is the main thing the CMS has to earn back against file-based MDX, and it directly serves the goal: showing judgement in context. Ship with one block (a callout/aside for "what this number doesn't tell you"); add a live-chart block once a post needs it.

Security

/admin is the first authenticated surface on this site. Public, hardened:

  • PAYLOAD_SECRET — long, random, set in the Portainer stack environment, never committed. The same variable must exist in staging with a different value.
  • Strong unique password on the single admin account.
  • Login rate limiting via Payload's maxLoginAttempts / lockTime.
  • X-Robots-Tag: noindex, nofollow on /admin/* and /cms-api/*, and a robots.ts disallow. The admin panel must never be indexed.
  • Public user creation disabled; no open registration.
  • Verify the existing CSP frame-ancestors directive does not break the admin panel.

Residual risk, accepted: a future Payload authentication CVE is live against the public internet. Mitigation is prompt patching, which the staging→prod pipeline already supports. If this becomes uncomfortable, restricting /admin at the proxy to LAN/VPN is a one-line change later.

Staging note: staging runs the same image on stx., so it gets its own admin panel and its own database. It must have its own PAYLOAD_SECRET and its own credentials — never production's.

Deployment changes

  • nextjs-app/Dockerfile — create and chown /app/media; ensure Payload's admin bundle and sharp survive standalone output file tracing.
  • docker-compose.portainer.yml and the staging equivalent — add DATABASE_URL and PAYLOAD_SECRET to the frontend service, add a payload_media volume mounted at /app/media, and add depends_on: sc_database.
  • Document both new environment variables in the compose header comment block, which is where this stack records its configuration.

SEO

  • Person (Tudor, with photo) and Organization JSON-LD on /about.
  • BlogPosting + BreadcrumbList on post pages, with author referencing the same Person.
  • Canonical URLs on /blog and every post.
  • Posts and /about added to the existing sitemap (app/sitemap.xml/route.ts and app/sitemaps/[...parts]). Post URLs come from Payload at request time.
  • RSS feed at /blog/rss.xml.
  • Footer links to both pages, under a new "About" column.

Navigation is deliberately left alone. Navigation.tsx renders a bottom tab bar on mobile that already carries four items (Search, Compare, Rankings, Admissions). A fifth tab makes each one cramped at 320px, and About and Blog are both lower-intent than any of the four. Both live in the footer; About additionally gets a byline link from every post, which is where a reader who cares actually asks the question. Revisit only if analytics show people hunting for it.

Testing

Unit (Jest):

  • Post rendering, including a post with no hero image and one with no excerpt.
  • Slug generation and collision handling.
  • JSON-LD shape for BlogPosting and Person.
  • The next.config.mjs conversion preserves the staging X-Robots-Tag rule — this guards the riskiest mechanical change in the plan.

E2E (Playwright, e2e/, required by CLAUDE.md for user-facing change):

  • /about renders, shows the author name and photo, and is reachable from the footer and nav.
  • /blog lists at least one post; clicking through reaches the post.
  • A post page renders title, date, body and byline.
  • /admin responds with noindex and does not leak a stack trace when unauthenticated.

Note the known constraint: new journeys cannot be proven in PR checks, because the staging E2E gate runs post-merge.

Risks

Risk Mitigation
next.config.mjs conversion silently drops the staging noindex header, making staging a crawlable duplicate Unit test asserting the header rule; verify on staging before promotion
Payload API route collides with the FastAPI /api proxy Remap to /cms-api; add to the proxy's exclusion list
Media volume mounts root-owned; all uploads fail with EACCES mkdir+chown in the Dockerfile before the mount; test an upload on staging
Build fails or bakes empty content because CI has no DB No build-time DB access; ISR only
A pipeline --drop destroys blog content Separate payload schema; verify --drop blast radius before building
Media volume not backed up; images unrecoverable Add payload_media to the backup routine
Payload auth CVE exposed publicly Prompt patching; proxy restriction available as a fallback
Blog launches empty or goes stale Ship with one post; cadence is explicitly "a few times a year", so no cadence is promised anywhere on the page — no dates implying a schedule

Sequence

Each step is independently reviewable and mergeable.

  1. Payload foundation — install, next.config.mjs conversion, payload schema, /cms-api remap, users collection, /admin hardening, compose and Dockerfile changes. No public-facing change yet. Verify on staging that the site is unchanged and /admin works.
  2. /about — coded page, photo, Person/Organization JSON-LD, footer and nav links, e2e journey. Independently valuable and does not depend on the blog.
  3. Blog — posts and media collections, /blog index and post pages, ISR plus revalidation hooks, RSS, sitemap, structured data, e2e journeys.
  4. First post — written in the admin panel, published through the normal flow, proving the whole path end to end.

Step 1 carries all the infrastructure risk and none of the visible benefit, so it should be verified on staging carefully before step 2 starts.

Dependencies on Tudor

  • A photograph. Blocks step 2. Nothing else in the plan is blocked by it.
  • The first post's subject. Blocks step 4 only. Suggested: what school performance data cannot tell you — it demonstrates judgement, is genuinely useful, and is the kind of thing an anonymous or machine-written site will not publish.
  • Confirmation that scripts/migrate_csv_to_db.py --drop is schema-scoped. Resolved 2026-09-02 — verified in backend/migration.py; see the Database section. No action needed.