diff --git a/docs/superpowers/specs/2026-09-02-about-and-blog-design.md b/docs/superpowers/specs/2026-09-02-about-and-blog-design.md new file mode 100644 index 0000000..3b2a99f --- /dev/null +++ b/docs/superpowers/specs/2026-09-02-about-and-blog-design.md @@ -0,0 +1,341 @@ +# Giving schoolcompare a human author: an About page and a blog + +**Date:** 2026-09-02 +**Status:** Design — awaiting review +**Scope:** A named author for the site, an `/about` page, and a Payload-CMS-backed +blog at `/blog`. + +## Why + +The site reads as synthetic. Not because of its tone, but because of three +specific absences: + +1. **Nobody is accountable for the numbers.** There is no author, no statement + of why the site exists, and no one who can be wrong. The only human trace on + the entire site is `contact@schoolcompare.co.uk` in the footer. +2. **No visible judgement.** Every figure is presented as though it fell out of + a machine. Hundreds of editorial decisions went into this codebase — which + metrics to show, when a benchmark is invalid, what to suppress — and not one + of them is visible to a reader. `isSpecialSchool()` silently drops the + England comparison for special schools and PRUs because that comparison is + meaningless; nowhere does the site *say* so. +3. **The voice is institutional third person.** "schoolcompare brings it all + into one place." "Built for parents, governors, journalists." That is + brochure register, and it is precisely the register that machine-generated + content defaults to. + +There is a second, independent reason. The SEO programme +(`2026-08-20-seo-programme-design.md`) defines eight workstreams and none of +them address E-E-A-T or authorship. School performance data is YMYL territory; +an anonymous site republishing DfE figures has no authorship signal at all. This +work fills that hole, and the blog gives W6 (explainer content) somewhere to +live. + +### The failure mode to avoid + +The standard fix — a stock photo and "Hi, I'm Tudor, and I'm passionate about +education!" — reads as *more* synthetic than the current coldness. Manufactured +warmth is a stronger machine-tell than plain institutional voice. Everything +here has to be specific, occasionally awkward, and willing to be unflattering, +or it makes the problem worse. + +## Positioning + +The author is **Tudor**: first name only, real photograph, no surname, no +employer named. + +The credibility claim is deliberately **not** educational expertise. The About +page states plainly: *"I'm not an education expert."* Authority comes from two +things that are actually true: + +- **Experience.** A parent going through primary admissions in south-west London + right now. Google's E-E-A-T leads with Experience, and lived experience of the + thing is exactly what the DfE's own service lacks. +- **Method.** Every number's provenance is stated, so a reader can check the + site rather than trust it. + +This is more durable than borrowed expertise: it cannot be undermined by someone +noticing the author has no teaching qualification. + +**Consequence for the design.** A `Person` entity with no surname is a weak +search signal and cannot be corroborated off-site. The credibility load +therefore shifts onto the methodology being visibly rigorous. That is a design +constraint, not a caveat — it is why the About page carries a substantial +"how this is built and where it can be wrong" section rather than a short bio. + +### Voice rules + +Applied to About and every post. Recorded here so the voice does not drift. + +- First person singular. "I built", not "we provide". +- Concrete over general. "when we were looking at schools in Wandsworth" beats + any amount of stated warmth. +- State limits before someone else finds them. Every post that presents a + metric says what it does not show. +- No mission statements, no "passionate about", no invented team. +- Short sentences. The existing code comments in this repo are already written + this way; the prose should match. + +## Scope + +**In:** + +- `/about` — a coded page (not CMS-managed). +- `/blog` and `/blog/[slug]` — Payload-backed, with an index and post pages. +- Payload CMS installed into the existing Next application. +- Footer and navigation links to both. +- `Person`, `Organization`, `BlogPosting`, `BreadcrumbList` JSON-LD. +- RSS feed and sitemap integration. +- One first post, so the blog does not launch empty. + +**Out (deliberately):** + +- Rewriting existing homepage/how-it-works copy into first person. Worth doing, + but it would double the review surface of this PR. Separate change. +- In-product signed notes on school pages (the "distributed humanity" idea). + Revisit once About and the blog exist. +- Comments, newsletter, author accounts beyond one. +- A team page. There is no team. + +## Architecture + +### Topology + +Payload 3 installs **into the existing Next application** and serves `/admin` +from the same container. One image, one deploy, no new service. This is +Payload 3's native model and it makes on-demand revalidation trivial, because +the CMS hooks run in the same process as the Next cache. + +Accepted costs: the public site's image now carries Payload, so a CMS security +patch redeploys the whole site; and the image grows substantially. + +### Two collisions that must be handled + +**1. `/api` is already taken.** `app/api/[...path]/route.ts` is a catch-all that +proxies `/api/*` to FastAPI at runtime. Payload's default API route is also +`/api`. Left alone, these fight, and the failure is not clean — the catch-all +would swallow Payload's admin API calls and forward them to FastAPI. + +Payload's API route is therefore remapped: + +```ts +routes: { api: '/cms-api', admin: '/admin' } +``` + +with its route group at `app/(payload)/cms-api/[...slug]/route.ts`. The +`/cms-api` prefix must also be added to the FastAPI proxy's excluded-paths list +as a defensive second line. + +**2. `next.config.js` is CommonJS.** Payload's `withPayload()` wrapper is ESM +only. The config must become `next.config.mjs`, converting `module.exports` to +`export default` and wrapping the export. All existing content — the standalone +output, `outputFileTracingIncludes`, the staging `X-Robots-Tag` header block, +the CSP — carries over unchanged. This is mechanical but it touches the file +that controls staging's noindex, so it needs care and an explicit test. + +### Database + +Payload uses the existing `sc_database` Postgres instance, in its **own +`payload` schema**: + +```ts +db: postgresAdapter({ + pool: { connectionString: process.env.DATABASE_URL }, + schemaName: 'payload', +}) +``` + +The frontend container is already on the `backend` Docker network, so it can +reach `sc_database:5432` with no networking change. It needs a new +`DATABASE_URL` environment variable. + +Schema isolation is not cosmetic. `public` currently holds the application +tables and Airflow's metadata, and `scripts/migrate_csv_to_db.py --drop` exists +to drop and reimport. Blog content living in its own schema means no data +pipeline operation can destroy it. **Before implementation, confirm that +`--drop` is schema-scoped and cannot reach `payload`.** + +Putting CMS tables in this instance is consistent with existing practice — +Airflow already stores its metadata there. + +### Migrations + +Payload's Postgres adapter auto-pushes schema in development and requires +explicit migrations in production. Use `prodMigrations`, which runs pending +migrations during server initialisation: + +```ts +db: postgresAdapter({ /* ... */, prodMigrations: migrations }) +``` + +This is preferred over a one-shot init container (the `airflow-init` pattern) +because the app is a single long-running process and there is no ordering +problem to solve. Migration files are generated with `payload migrate:create` +and committed, so schema changes travel through the same PR and staging gate as +code. + +### Media + +Uploads go to a Docker named volume, consistent with `postgres_data`, +`typesense_data` and `airflow_logs`. + +- `staticDir` must be an **absolute** path in Payload 3: `/app/media`. +- The container runs as `nextjs` (uid 1001). The Dockerfile must + `mkdir -p /app/media && chown nextjs:nodejs /app/media` **before** the volume + is mounted, or Docker will create the mountpoint root-owned and every upload + will fail with EACCES. +- `sharp` moves from `devDependencies` to `dependencies` — Payload needs it at + runtime to generate `imageSizes`. +- The volume must be added to the backup routine alongside Postgres. A blog + post's images are not reproducible from the pipeline. + +### Rendering + +**Constraint:** CI builds the image with no database reachable. Blog pages +therefore cannot use build-time `generateStaticParams` — that would either fail +the build or bake in an empty post list. + +Instead: ISR. Post and index pages declare a `revalidate` window and render on +first request, with Payload `afterChange` / `afterDelete` hooks calling +`revalidatePath('/blog')` and `revalidatePath('/blog/' + slug)` for immediate +publication. Because Payload runs in the same process, the hook calls +`revalidatePath` from `next/cache` directly — no webhook, no shared secret. + +The ISR cache lives on container disk and is cleared by a redeploy. For a +single container serving a handful of posts this is fine. + +### Collections + +- **`posts`** — `title`, `slug`, `publishedAt`, `excerpt`, `heroImage` + (relation to `media`), `content` (Lexical rich text), `seo` group + (`metaTitle`, `metaDescription`), `_status` (drafts enabled). +- **`media`** — upload collection, `alt` required, `imageSizes` for thumbnail + and hero widths, public read access. +- **`users`** — Payload's auth collection. One account. Public creation + disabled. + +Drafts are enabled so posts can be written over several sittings and previewed +before publication. + +**Payload Blocks** are how posts embed live product components — a real trend +chart or comparison table inside a post, rendered from live data rather than +screenshotted. This is the main thing the CMS has to earn back against +file-based MDX, and it directly serves the goal: showing judgement in context. +Ship with one block (a callout/aside for "what this number doesn't tell you"); +add a live-chart block once a post needs it. + +### Security + +`/admin` is the first authenticated surface on this site. Public, hardened: + +- `PAYLOAD_SECRET` — long, random, set in the Portainer stack environment, never + committed. The same variable must exist in staging with a *different* value. +- Strong unique password on the single admin account. +- Login rate limiting via Payload's `maxLoginAttempts` / `lockTime`. +- `X-Robots-Tag: noindex, nofollow` on `/admin/*` and `/cms-api/*`, and a + `robots.ts` disallow. The admin panel must never be indexed. +- Public user creation disabled; no open registration. +- Verify the existing CSP `frame-ancestors` directive does not break the admin + panel. + +Residual risk, accepted: a future Payload authentication CVE is live against the +public internet. Mitigation is prompt patching, which the staging→prod pipeline +already supports. If this becomes uncomfortable, restricting `/admin` at the +proxy to LAN/VPN is a one-line change later. + +Staging note: staging runs the same image on `stx.`, so it gets its own admin +panel and its own database. It must have its own `PAYLOAD_SECRET` and its own +credentials — never production's. + +## Deployment changes + +- `nextjs-app/Dockerfile` — create and chown `/app/media`; ensure Payload's + admin bundle and `sharp` survive standalone output file tracing. +- `docker-compose.portainer.yml` and the staging equivalent — add + `DATABASE_URL` and `PAYLOAD_SECRET` to the `frontend` service, add a + `payload_media` volume mounted at `/app/media`, and add + `depends_on: sc_database`. +- Document both new environment variables in the compose header comment block, + which is where this stack records its configuration. + +## SEO + +- `Person` (Tudor, with photo) and `Organization` JSON-LD on `/about`. +- `BlogPosting` + `BreadcrumbList` on post pages, with `author` referencing the + same `Person`. +- Canonical URLs on `/blog` and every post. +- Posts and `/about` added to the existing sitemap (`app/sitemap.xml/route.ts` + and `app/sitemaps/[...parts]`). Post URLs come from Payload at request time. +- RSS feed at `/blog/rss.xml`. +- Footer links to both pages, under a new "About" column. + +**Navigation is deliberately left alone.** `Navigation.tsx` renders a bottom tab +bar on mobile that already carries four items (Search, Compare, Rankings, +Admissions). A fifth tab makes each one cramped at 320px, and About and Blog are +both lower-intent than any of the four. Both live in the footer; About +additionally gets a byline link from every post, which is where a reader who +cares actually asks the question. Revisit only if analytics show people hunting +for it. + +## Testing + +Unit (Jest): + +- Post rendering, including a post with no hero image and one with no excerpt. +- Slug generation and collision handling. +- JSON-LD shape for `BlogPosting` and `Person`. +- The `next.config.mjs` conversion preserves the staging `X-Robots-Tag` rule — + this guards the riskiest mechanical change in the plan. + +E2E (Playwright, `e2e/`, required by CLAUDE.md for user-facing change): + +- `/about` renders, shows the author name and photo, and is reachable from the + footer and nav. +- `/blog` lists at least one post; clicking through reaches the post. +- A post page renders title, date, body and byline. +- `/admin` responds with `noindex` and does not leak a stack trace when + unauthenticated. + +Note the known constraint: new journeys cannot be proven in PR checks, because +the staging E2E gate runs post-merge. + +## Risks + +| Risk | Mitigation | +|---|---| +| `next.config.mjs` conversion silently drops the staging noindex header, making staging a crawlable duplicate | Unit test asserting the header rule; verify on staging before promotion | +| Payload API route collides with the FastAPI `/api` proxy | Remap to `/cms-api`; add to the proxy's exclusion list | +| Media volume mounts root-owned; all uploads fail with EACCES | `mkdir`+`chown` in the Dockerfile before the mount; test an upload on staging | +| Build fails or bakes empty content because CI has no DB | No build-time DB access; ISR only | +| A pipeline `--drop` destroys blog content | Separate `payload` schema; verify `--drop` blast radius before building | +| Media volume not backed up; images unrecoverable | Add `payload_media` to the backup routine | +| Payload auth CVE exposed publicly | Prompt patching; proxy restriction available as a fallback | +| Blog launches empty or goes stale | Ship with one post; cadence is explicitly "a few times a year", so no cadence is promised anywhere on the page — no dates implying a schedule | + +## Sequence + +Each step is independently reviewable and mergeable. + +1. **Payload foundation** — install, `next.config.mjs` conversion, `payload` + schema, `/cms-api` remap, `users` collection, `/admin` hardening, compose and + Dockerfile changes. No public-facing change yet. Verify on staging that the + site is unchanged and `/admin` works. +2. **`/about`** — coded page, photo, `Person`/`Organization` JSON-LD, footer and + nav links, e2e journey. Independently valuable and does not depend on the + blog. +3. **Blog** — `posts` and `media` collections, `/blog` index and post pages, ISR + plus revalidation hooks, RSS, sitemap, structured data, e2e journeys. +4. **First post** — written in the admin panel, published through the normal + flow, proving the whole path end to end. + +Step 1 carries all the infrastructure risk and none of the visible benefit, so +it should be verified on staging carefully before step 2 starts. + +## Dependencies on Tudor + +- **A photograph.** Blocks step 2. Nothing else in the plan is blocked by it. +- **The first post's subject.** Blocks step 4 only. Suggested: what school + performance data cannot tell you — it demonstrates judgement, is genuinely + useful, and is the kind of thing an anonymous or machine-written site will not + publish. +- **Confirmation** that `scripts/migrate_csv_to_db.py --drop` is schema-scoped.