# Giving schoolcompare a human author: an About page and a blog **Date:** 2026-09-02 **Status:** Design — awaiting review **Scope:** A named author for the site, an `/about` page, and a Payload-CMS-backed blog at `/blog`. ## Why The site reads as synthetic. Not because of its tone, but because of three specific absences: 1. **Nobody is accountable for the numbers.** There is no author, no statement of why the site exists, and no one who can be wrong. The only human trace on the entire site is `contact@schoolcompare.co.uk` in the footer. 2. **No visible judgement.** Every figure is presented as though it fell out of a machine. Hundreds of editorial decisions went into this codebase — which metrics to show, when a benchmark is invalid, what to suppress — and not one of them is visible to a reader. `isSpecialSchool()` silently drops the England comparison for special schools and PRUs because that comparison is meaningless; nowhere does the site *say* so. 3. **The voice is institutional third person.** "schoolcompare brings it all into one place." "Built for parents, governors, journalists." That is brochure register, and it is precisely the register that machine-generated content defaults to. There is a second, independent reason. The SEO programme (`2026-08-20-seo-programme-design.md`) defines eight workstreams and none of them address E-E-A-T or authorship. School performance data is YMYL territory; an anonymous site republishing DfE figures has no authorship signal at all. This work fills that hole, and the blog gives W6 (explainer content) somewhere to live. ### The failure mode to avoid The standard fix — a stock photo and "Hi, I'm Tudor, and I'm passionate about education!" — reads as *more* synthetic than the current coldness. Manufactured warmth is a stronger machine-tell than plain institutional voice. Everything here has to be specific, occasionally awkward, and willing to be unflattering, or it makes the problem worse. ## Positioning The author is **Tudor**: first name only, real photograph, no surname, no employer named. The credibility claim is deliberately **not** educational expertise. The About page states plainly: *"I'm not an education expert."* Authority comes from two things that are actually true: - **Experience.** A parent going through primary admissions in south-west London right now. Google's E-E-A-T leads with Experience, and lived experience of the thing is exactly what the DfE's own service lacks. - **Method.** Every number's provenance is stated, so a reader can check the site rather than trust it. This is more durable than borrowed expertise: it cannot be undermined by someone noticing the author has no teaching qualification. **Consequence for the design.** A `Person` entity with no surname is a weak search signal and cannot be corroborated off-site. The credibility load therefore shifts onto the methodology being visibly rigorous. That is a design constraint, not a caveat — it is why the About page carries a substantial "how this is built and where it can be wrong" section rather than a short bio. ### Voice rules Applied to About and every post. Recorded here so the voice does not drift. - First person singular. "I built", not "we provide". - Concrete over general. "when we were looking at schools in Wandsworth" beats any amount of stated warmth. - State limits before someone else finds them. Every post that presents a metric says what it does not show. - No mission statements, no "passionate about", no invented team. - No em dashes. One of the clearest tells of machine-written prose, which is the exact problem this work exists to fix. - Short sentences. The existing code comments in this repo are already written this way; the prose should match. ## Scope **In:** - `/about` — a coded page (not CMS-managed). - `/blog` and `/blog/[slug]` — Payload-backed, with an index and post pages. - Payload CMS installed into the existing Next application. - Footer and navigation links to both. - `Person`, `Organization`, `BlogPosting`, `BreadcrumbList` JSON-LD. - RSS feed and sitemap integration. - One first post, so the blog does not launch empty. **Out (deliberately):** - Rewriting existing homepage/how-it-works copy into first person. Worth doing, but it would double the review surface of this PR. Separate change. - In-product signed notes on school pages (the "distributed humanity" idea). Revisit once About and the blog exist. - Comments, newsletter, author accounts beyond one. - A team page. There is no team. ## Architecture ### Topology Payload 3 installs **into the existing Next application** and serves `/admin` from the same container. One image, one deploy, no new service. This is Payload 3's native model and it makes on-demand revalidation trivial, because the CMS hooks run in the same process as the Next cache. Accepted costs: the public site's image now carries Payload, so a CMS security patch redeploys the whole site; and the image grows substantially. ### Two collisions that must be handled **1. `/api` is already taken.** `app/api/[...path]/route.ts` is a catch-all that proxies `/api/*` to FastAPI at runtime. Payload's default API route is also `/api`. Left alone, these fight, and the failure is not clean — the catch-all would swallow Payload's admin API calls and forward them to FastAPI. Payload's API route is therefore remapped: ```ts routes: { api: '/cms-api', admin: '/admin' } ``` with its route group at `app/(payload)/cms-api/[...slug]/route.ts`. The `/cms-api` prefix must also be added to the FastAPI proxy's excluded-paths list as a defensive second line. **2. `next.config.js` is CommonJS.** Payload's `withPayload()` wrapper is ESM only. The config must become `next.config.mjs`, converting `module.exports` to `export default` and wrapping the export. All existing content — the standalone output, `outputFileTracingIncludes`, the staging `X-Robots-Tag` header block, the CSP — carries over unchanged. This is mechanical but it touches the file that controls staging's noindex, so it needs care and an explicit test. ### Database Payload uses the existing `sc_database` Postgres instance, in its **own `payload` schema**: ```ts db: postgresAdapter({ pool: { connectionString: process.env.DATABASE_URL }, schemaName: 'payload', }) ``` The frontend container is already on the `backend` Docker network, so it can reach `sc_database:5432` with no networking change. It needs a new `DATABASE_URL` environment variable. Schema isolation is not cosmetic. `public` currently holds the application tables and Airflow's metadata, and `scripts/migrate_csv_to_db.py --drop` exists to drop and reimport. Blog content living in its own schema means no data pipeline operation can destroy it. **Verified 2026-09-02** (this was an open question when the spec was written). `--drop` calls `run_full_migration()` in `backend/migration.py`, which drops exactly two tables by name: ```python ks2_tables = ["school_results", "schools"] for tname in ks2_tables: if tname in existing: Base.metadata.tables[tname].drop(bind=engine) ``` There is no `Base.metadata.drop_all()` anywhere in `backend/`, and no `DROP SCHEMA`. The only other drop is `_apply_schema_drops()`, a single schema-qualified `DROP TABLE IF EXISTS marts.fact_parent_view CASCADE`. Nothing sets `search_path`, so the SQLAlchemy metadata resolves to `public`, and `inspector.get_table_names()` does not even enumerate other schemas. So the guarantee is stronger than schema isolation alone: `--drop` targets two named tables that Payload does not have, and would not reach `posts`, `media` or `users` even if they shared a schema. The `payload` schema remains the right choice — it protects against a *future* broadening of that script rather than today's behaviour — but the safety claim rests on verified code, not on assumption. Putting CMS tables in this instance is consistent with existing practice — Airflow already stores its metadata there. ### Migrations Payload's Postgres adapter auto-pushes schema in development and requires explicit migrations in production. Use `prodMigrations`, which runs pending migrations during server initialisation: ```ts db: postgresAdapter({ /* ... */, prodMigrations: migrations }) ``` This is preferred over a one-shot init container (the `airflow-init` pattern) because the app is a single long-running process and there is no ordering problem to solve. Migration files are generated with `payload migrate:create` and committed, so schema changes travel through the same PR and staging gate as code. ### Media Uploads go to a Docker named volume, consistent with `postgres_data`, `typesense_data` and `airflow_logs`. - `staticDir` must be an **absolute** path in Payload 3: `/app/media`. - The container runs as `nextjs` (uid 1001). The Dockerfile must `mkdir -p /app/media && chown nextjs:nodejs /app/media` **before** the volume is mounted, or Docker will create the mountpoint root-owned and every upload will fail with EACCES. - `sharp` moves from `devDependencies` to `dependencies` — Payload needs it at runtime to generate `imageSizes`. - The volume must be added to the backup routine alongside Postgres. A blog post's images are not reproducible from the pipeline. ### Rendering **Constraint:** CI builds the image with no database reachable. Blog pages therefore cannot use build-time `generateStaticParams` — that would either fail the build or bake in an empty post list. Instead: ISR. Post and index pages declare a `revalidate` window and render on first request, with Payload `afterChange` / `afterDelete` hooks calling `revalidatePath('/blog')` and `revalidatePath('/blog/' + slug)` for immediate publication. Because Payload runs in the same process, the hook calls `revalidatePath` from `next/cache` directly — no webhook, no shared secret. The ISR cache lives on container disk and is cleared by a redeploy. For a single container serving a handful of posts this is fine. ### Collections - **`posts`** — `title`, `slug`, `publishedAt`, `excerpt`, `heroImage` (relation to `media`), `content` (Lexical rich text), `seo` group (`metaTitle`, `metaDescription`), `_status` (drafts enabled). - **`media`** — upload collection, `alt` required, `imageSizes` for thumbnail and hero widths, public read access. - **`users`** — Payload's auth collection. One account. Public creation disabled. Drafts are enabled so posts can be written over several sittings and previewed before publication. **Payload Blocks** are how posts embed live product components — a real trend chart or comparison table inside a post, rendered from live data rather than screenshotted. This is the main thing the CMS has to earn back against file-based MDX, and it directly serves the goal: showing judgement in context. Ship with one block (a callout/aside for "what this number doesn't tell you"); add a live-chart block once a post needs it. ### Security `/admin` is the first authenticated surface on this site. Public, hardened: - `PAYLOAD_SECRET` — long, random, set in the Portainer stack environment, never committed. The same variable must exist in staging with a *different* value. - Strong unique password on the single admin account. - Login rate limiting via Payload's `maxLoginAttempts` / `lockTime`. - `X-Robots-Tag: noindex, nofollow` on `/admin/*` and `/cms-api/*`, and a `robots.ts` disallow. The admin panel must never be indexed. - Public user creation disabled; no open registration. - Verify the existing CSP `frame-ancestors` directive does not break the admin panel. Residual risk, accepted: a future Payload authentication CVE is live against the public internet. Mitigation is prompt patching, which the staging→prod pipeline already supports. If this becomes uncomfortable, restricting `/admin` at the proxy to LAN/VPN is a one-line change later. Staging note: staging runs the same image on `stx.`, so it gets its own admin panel and its own database. It must have its own `PAYLOAD_SECRET` and its own credentials — never production's. ## Deployment changes - `nextjs-app/Dockerfile` — create and chown `/app/media`; ensure Payload's admin bundle and `sharp` survive standalone output file tracing. - `docker-compose.portainer.yml` and the staging equivalent — add `DATABASE_URL` and `PAYLOAD_SECRET` to the `frontend` service, add a `payload_media` volume mounted at `/app/media`, and add `depends_on: sc_database`. - Document both new environment variables in the compose header comment block, which is where this stack records its configuration. ## SEO - `Person` (Tudor, with photo) and `Organization` JSON-LD on `/about`. - `BlogPosting` + `BreadcrumbList` on post pages, with `author` referencing the same `Person`. - Canonical URLs on `/blog` and every post. - Posts and `/about` added to the existing sitemap (`app/sitemap.xml/route.ts` and `app/sitemaps/[...parts]`). Post URLs come from Payload at request time. - RSS feed at `/blog/rss.xml`. - Footer links to both pages, under a new "About" column. **Navigation is deliberately left alone.** `Navigation.tsx` renders a bottom tab bar on mobile that already carries four items (Search, Compare, Rankings, Admissions). A fifth tab makes each one cramped at 320px, and About and Blog are both lower-intent than any of the four. Both live in the footer; About additionally gets a byline link from every post, which is where a reader who cares actually asks the question. Revisit only if analytics show people hunting for it. ## Testing Unit (Jest): - Post rendering, including a post with no hero image and one with no excerpt. - Slug generation and collision handling. - JSON-LD shape for `BlogPosting` and `Person`. - The `next.config.mjs` conversion preserves the staging `X-Robots-Tag` rule — this guards the riskiest mechanical change in the plan. E2E (Playwright, `e2e/`, required by CLAUDE.md for user-facing change): - `/about` renders, shows the author name and photo, and is reachable from the footer and nav. - `/blog` lists at least one post; clicking through reaches the post. - A post page renders title, date, body and byline. - `/admin` responds with `noindex` and does not leak a stack trace when unauthenticated. Note the known constraint: new journeys cannot be proven in PR checks, because the staging E2E gate runs post-merge. ## Risks | Risk | Mitigation | |---|---| | `next.config.mjs` conversion silently drops the staging noindex header, making staging a crawlable duplicate | Unit test asserting the header rule; verify on staging before promotion | | Payload API route collides with the FastAPI `/api` proxy | Remap to `/cms-api`; add to the proxy's exclusion list | | Media volume mounts root-owned; all uploads fail with EACCES | `mkdir`+`chown` in the Dockerfile before the mount; test an upload on staging | | Build fails or bakes empty content because CI has no DB | No build-time DB access; ISR only | | A pipeline `--drop` destroys blog content | Separate `payload` schema; verify `--drop` blast radius before building | | Media volume not backed up; images unrecoverable | Add `payload_media` to the backup routine | | Payload auth CVE exposed publicly | Prompt patching; proxy restriction available as a fallback | | Blog launches empty or goes stale | Ship with one post; cadence is explicitly "a few times a year", so no cadence is promised anywhere on the page — no dates implying a schedule | ## Sequence Each step is independently reviewable and mergeable. 1. **Payload foundation** — install, `next.config.mjs` conversion, `payload` schema, `/cms-api` remap, `users` collection, `/admin` hardening, compose and Dockerfile changes. No public-facing change yet. Verify on staging that the site is unchanged and `/admin` works. 2. **`/about`** — coded page, photo, `Person`/`Organization` JSON-LD, footer and nav links, e2e journey. Independently valuable and does not depend on the blog. 3. **Blog** — `posts` and `media` collections, `/blog` index and post pages, ISR plus revalidation hooks, RSS, sitemap, structured data, e2e journeys. 4. **First post** — written in the admin panel, published through the normal flow, proving the whole path end to end. Step 1 carries all the infrastructure risk and none of the visible benefit, so it should be verified on staging carefully before step 2 starts. ## Dependencies on Tudor - **A photograph.** Blocks step 2. Nothing else in the plan is blocked by it. - **The first post's subject.** Blocks step 4 only. Suggested: what school performance data cannot tell you — it demonstrates judgement, is genuinely useful, and is the kind of thing an anonymous or machine-written site will not publish. - ~~Confirmation that `scripts/migrate_csv_to_db.py --drop` is schema-scoped.~~ **Resolved 2026-09-02** — verified in `backend/migration.py`; see the Database section. No action needed.