The em dash is one of the clearest tells of machine-written text, which is the exact impression this work exists to remove. Rewritten rather than substituted: where a dash was carrying a real aside the sentence is split or recast, not patched with a comma. Covers the About page, the two Callout labels an editor sees in the admin panel, and PUBLISHING.md, which defines the house style and should follow it. The rule is now recorded in that house style and in the spec's voice rules, so it survives this branch. Code comments are left alone: they are not copy, and the surrounding codebase uses the same punctuation throughout. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017YmbBhr8s7GusjDE12hrZM
369 lines
17 KiB
Markdown
369 lines
17 KiB
Markdown
# Giving schoolcompare a human author: an About page and a blog
|
|
|
|
**Date:** 2026-09-02
|
|
**Status:** Design — awaiting review
|
|
**Scope:** A named author for the site, an `/about` page, and a Payload-CMS-backed
|
|
blog at `/blog`.
|
|
|
|
## Why
|
|
|
|
The site reads as synthetic. Not because of its tone, but because of three
|
|
specific absences:
|
|
|
|
1. **Nobody is accountable for the numbers.** There is no author, no statement
|
|
of why the site exists, and no one who can be wrong. The only human trace on
|
|
the entire site is `contact@schoolcompare.co.uk` in the footer.
|
|
2. **No visible judgement.** Every figure is presented as though it fell out of
|
|
a machine. Hundreds of editorial decisions went into this codebase — which
|
|
metrics to show, when a benchmark is invalid, what to suppress — and not one
|
|
of them is visible to a reader. `isSpecialSchool()` silently drops the
|
|
England comparison for special schools and PRUs because that comparison is
|
|
meaningless; nowhere does the site *say* so.
|
|
3. **The voice is institutional third person.** "schoolcompare brings it all
|
|
into one place." "Built for parents, governors, journalists." That is
|
|
brochure register, and it is precisely the register that machine-generated
|
|
content defaults to.
|
|
|
|
There is a second, independent reason. The SEO programme
|
|
(`2026-08-20-seo-programme-design.md`) defines eight workstreams and none of
|
|
them address E-E-A-T or authorship. School performance data is YMYL territory;
|
|
an anonymous site republishing DfE figures has no authorship signal at all. This
|
|
work fills that hole, and the blog gives W6 (explainer content) somewhere to
|
|
live.
|
|
|
|
### The failure mode to avoid
|
|
|
|
The standard fix — a stock photo and "Hi, I'm Tudor, and I'm passionate about
|
|
education!" — reads as *more* synthetic than the current coldness. Manufactured
|
|
warmth is a stronger machine-tell than plain institutional voice. Everything
|
|
here has to be specific, occasionally awkward, and willing to be unflattering,
|
|
or it makes the problem worse.
|
|
|
|
## Positioning
|
|
|
|
The author is **Tudor**: first name only, real photograph, no surname, no
|
|
employer named.
|
|
|
|
The credibility claim is deliberately **not** educational expertise. The About
|
|
page states plainly: *"I'm not an education expert."* Authority comes from two
|
|
things that are actually true:
|
|
|
|
- **Experience.** A parent going through primary admissions in south-west London
|
|
right now. Google's E-E-A-T leads with Experience, and lived experience of the
|
|
thing is exactly what the DfE's own service lacks.
|
|
- **Method.** Every number's provenance is stated, so a reader can check the
|
|
site rather than trust it.
|
|
|
|
This is more durable than borrowed expertise: it cannot be undermined by someone
|
|
noticing the author has no teaching qualification.
|
|
|
|
**Consequence for the design.** A `Person` entity with no surname is a weak
|
|
search signal and cannot be corroborated off-site. The credibility load
|
|
therefore shifts onto the methodology being visibly rigorous. That is a design
|
|
constraint, not a caveat — it is why the About page carries a substantial
|
|
"how this is built and where it can be wrong" section rather than a short bio.
|
|
|
|
### Voice rules
|
|
|
|
Applied to About and every post. Recorded here so the voice does not drift.
|
|
|
|
- First person singular. "I built", not "we provide".
|
|
- Concrete over general. "when we were looking at schools in Wandsworth" beats
|
|
any amount of stated warmth.
|
|
- State limits before someone else finds them. Every post that presents a
|
|
metric says what it does not show.
|
|
- No mission statements, no "passionate about", no invented team.
|
|
- No em dashes. One of the clearest tells of machine-written prose, which is
|
|
the exact problem this work exists to fix.
|
|
- Short sentences. The existing code comments in this repo are already written
|
|
this way; the prose should match.
|
|
|
|
## Scope
|
|
|
|
**In:**
|
|
|
|
- `/about` — a coded page (not CMS-managed).
|
|
- `/blog` and `/blog/[slug]` — Payload-backed, with an index and post pages.
|
|
- Payload CMS installed into the existing Next application.
|
|
- Footer and navigation links to both.
|
|
- `Person`, `Organization`, `BlogPosting`, `BreadcrumbList` JSON-LD.
|
|
- RSS feed and sitemap integration.
|
|
- One first post, so the blog does not launch empty.
|
|
|
|
**Out (deliberately):**
|
|
|
|
- Rewriting existing homepage/how-it-works copy into first person. Worth doing,
|
|
but it would double the review surface of this PR. Separate change.
|
|
- In-product signed notes on school pages (the "distributed humanity" idea).
|
|
Revisit once About and the blog exist.
|
|
- Comments, newsletter, author accounts beyond one.
|
|
- A team page. There is no team.
|
|
|
|
## Architecture
|
|
|
|
### Topology
|
|
|
|
Payload 3 installs **into the existing Next application** and serves `/admin`
|
|
from the same container. One image, one deploy, no new service. This is
|
|
Payload 3's native model and it makes on-demand revalidation trivial, because
|
|
the CMS hooks run in the same process as the Next cache.
|
|
|
|
Accepted costs: the public site's image now carries Payload, so a CMS security
|
|
patch redeploys the whole site; and the image grows substantially.
|
|
|
|
### Two collisions that must be handled
|
|
|
|
**1. `/api` is already taken.** `app/api/[...path]/route.ts` is a catch-all that
|
|
proxies `/api/*` to FastAPI at runtime. Payload's default API route is also
|
|
`/api`. Left alone, these fight, and the failure is not clean — the catch-all
|
|
would swallow Payload's admin API calls and forward them to FastAPI.
|
|
|
|
Payload's API route is therefore remapped:
|
|
|
|
```ts
|
|
routes: { api: '/cms-api', admin: '/admin' }
|
|
```
|
|
|
|
with its route group at `app/(payload)/cms-api/[...slug]/route.ts`. The
|
|
`/cms-api` prefix must also be added to the FastAPI proxy's excluded-paths list
|
|
as a defensive second line.
|
|
|
|
**2. `next.config.js` is CommonJS.** Payload's `withPayload()` wrapper is ESM
|
|
only. The config must become `next.config.mjs`, converting `module.exports` to
|
|
`export default` and wrapping the export. All existing content — the standalone
|
|
output, `outputFileTracingIncludes`, the staging `X-Robots-Tag` header block,
|
|
the CSP — carries over unchanged. This is mechanical but it touches the file
|
|
that controls staging's noindex, so it needs care and an explicit test.
|
|
|
|
### Database
|
|
|
|
Payload uses the existing `sc_database` Postgres instance, in its **own
|
|
`payload` schema**:
|
|
|
|
```ts
|
|
db: postgresAdapter({
|
|
pool: { connectionString: process.env.DATABASE_URL },
|
|
schemaName: 'payload',
|
|
})
|
|
```
|
|
|
|
The frontend container is already on the `backend` Docker network, so it can
|
|
reach `sc_database:5432` with no networking change. It needs a new
|
|
`DATABASE_URL` environment variable.
|
|
|
|
Schema isolation is not cosmetic. `public` currently holds the application
|
|
tables and Airflow's metadata, and `scripts/migrate_csv_to_db.py --drop` exists
|
|
to drop and reimport. Blog content living in its own schema means no data
|
|
pipeline operation can destroy it.
|
|
|
|
**Verified 2026-09-02** (this was an open question when the spec was written).
|
|
`--drop` calls `run_full_migration()` in `backend/migration.py`, which drops
|
|
exactly two tables by name:
|
|
|
|
```python
|
|
ks2_tables = ["school_results", "schools"]
|
|
for tname in ks2_tables:
|
|
if tname in existing:
|
|
Base.metadata.tables[tname].drop(bind=engine)
|
|
```
|
|
|
|
There is no `Base.metadata.drop_all()` anywhere in `backend/`, and no
|
|
`DROP SCHEMA`. The only other drop is `_apply_schema_drops()`, a single
|
|
schema-qualified `DROP TABLE IF EXISTS marts.fact_parent_view CASCADE`.
|
|
Nothing sets `search_path`, so the SQLAlchemy metadata resolves to `public`,
|
|
and `inspector.get_table_names()` does not even enumerate other schemas.
|
|
|
|
So the guarantee is stronger than schema isolation alone: `--drop` targets two
|
|
named tables that Payload does not have, and would not reach `posts`, `media`
|
|
or `users` even if they shared a schema. The `payload` schema remains the right
|
|
choice — it protects against a *future* broadening of that script rather than
|
|
today's behaviour — but the safety claim rests on verified code, not on
|
|
assumption.
|
|
|
|
Putting CMS tables in this instance is consistent with existing practice —
|
|
Airflow already stores its metadata there.
|
|
|
|
### Migrations
|
|
|
|
Payload's Postgres adapter auto-pushes schema in development and requires
|
|
explicit migrations in production. Use `prodMigrations`, which runs pending
|
|
migrations during server initialisation:
|
|
|
|
```ts
|
|
db: postgresAdapter({ /* ... */, prodMigrations: migrations })
|
|
```
|
|
|
|
This is preferred over a one-shot init container (the `airflow-init` pattern)
|
|
because the app is a single long-running process and there is no ordering
|
|
problem to solve. Migration files are generated with `payload migrate:create`
|
|
and committed, so schema changes travel through the same PR and staging gate as
|
|
code.
|
|
|
|
### Media
|
|
|
|
Uploads go to a Docker named volume, consistent with `postgres_data`,
|
|
`typesense_data` and `airflow_logs`.
|
|
|
|
- `staticDir` must be an **absolute** path in Payload 3: `/app/media`.
|
|
- The container runs as `nextjs` (uid 1001). The Dockerfile must
|
|
`mkdir -p /app/media && chown nextjs:nodejs /app/media` **before** the volume
|
|
is mounted, or Docker will create the mountpoint root-owned and every upload
|
|
will fail with EACCES.
|
|
- `sharp` moves from `devDependencies` to `dependencies` — Payload needs it at
|
|
runtime to generate `imageSizes`.
|
|
- The volume must be added to the backup routine alongside Postgres. A blog
|
|
post's images are not reproducible from the pipeline.
|
|
|
|
### Rendering
|
|
|
|
**Constraint:** CI builds the image with no database reachable. Blog pages
|
|
therefore cannot use build-time `generateStaticParams` — that would either fail
|
|
the build or bake in an empty post list.
|
|
|
|
Instead: ISR. Post and index pages declare a `revalidate` window and render on
|
|
first request, with Payload `afterChange` / `afterDelete` hooks calling
|
|
`revalidatePath('/blog')` and `revalidatePath('/blog/' + slug)` for immediate
|
|
publication. Because Payload runs in the same process, the hook calls
|
|
`revalidatePath` from `next/cache` directly — no webhook, no shared secret.
|
|
|
|
The ISR cache lives on container disk and is cleared by a redeploy. For a
|
|
single container serving a handful of posts this is fine.
|
|
|
|
### Collections
|
|
|
|
- **`posts`** — `title`, `slug`, `publishedAt`, `excerpt`, `heroImage`
|
|
(relation to `media`), `content` (Lexical rich text), `seo` group
|
|
(`metaTitle`, `metaDescription`), `_status` (drafts enabled).
|
|
- **`media`** — upload collection, `alt` required, `imageSizes` for thumbnail
|
|
and hero widths, public read access.
|
|
- **`users`** — Payload's auth collection. One account. Public creation
|
|
disabled.
|
|
|
|
Drafts are enabled so posts can be written over several sittings and previewed
|
|
before publication.
|
|
|
|
**Payload Blocks** are how posts embed live product components — a real trend
|
|
chart or comparison table inside a post, rendered from live data rather than
|
|
screenshotted. This is the main thing the CMS has to earn back against
|
|
file-based MDX, and it directly serves the goal: showing judgement in context.
|
|
Ship with one block (a callout/aside for "what this number doesn't tell you");
|
|
add a live-chart block once a post needs it.
|
|
|
|
### Security
|
|
|
|
`/admin` is the first authenticated surface on this site. Public, hardened:
|
|
|
|
- `PAYLOAD_SECRET` — long, random, set in the Portainer stack environment, never
|
|
committed. The same variable must exist in staging with a *different* value.
|
|
- Strong unique password on the single admin account.
|
|
- Login rate limiting via Payload's `maxLoginAttempts` / `lockTime`.
|
|
- `X-Robots-Tag: noindex, nofollow` on `/admin/*` and `/cms-api/*`, and a
|
|
`robots.ts` disallow. The admin panel must never be indexed.
|
|
- Public user creation disabled; no open registration.
|
|
- Verify the existing CSP `frame-ancestors` directive does not break the admin
|
|
panel.
|
|
|
|
Residual risk, accepted: a future Payload authentication CVE is live against the
|
|
public internet. Mitigation is prompt patching, which the staging→prod pipeline
|
|
already supports. If this becomes uncomfortable, restricting `/admin` at the
|
|
proxy to LAN/VPN is a one-line change later.
|
|
|
|
Staging note: staging runs the same image on `stx.`, so it gets its own admin
|
|
panel and its own database. It must have its own `PAYLOAD_SECRET` and its own
|
|
credentials — never production's.
|
|
|
|
## Deployment changes
|
|
|
|
- `nextjs-app/Dockerfile` — create and chown `/app/media`; ensure Payload's
|
|
admin bundle and `sharp` survive standalone output file tracing.
|
|
- `docker-compose.portainer.yml` and the staging equivalent — add
|
|
`DATABASE_URL` and `PAYLOAD_SECRET` to the `frontend` service, add a
|
|
`payload_media` volume mounted at `/app/media`, and add
|
|
`depends_on: sc_database`.
|
|
- Document both new environment variables in the compose header comment block,
|
|
which is where this stack records its configuration.
|
|
|
|
## SEO
|
|
|
|
- `Person` (Tudor, with photo) and `Organization` JSON-LD on `/about`.
|
|
- `BlogPosting` + `BreadcrumbList` on post pages, with `author` referencing the
|
|
same `Person`.
|
|
- Canonical URLs on `/blog` and every post.
|
|
- Posts and `/about` added to the existing sitemap (`app/sitemap.xml/route.ts`
|
|
and `app/sitemaps/[...parts]`). Post URLs come from Payload at request time.
|
|
- RSS feed at `/blog/rss.xml`.
|
|
- Footer links to both pages, under a new "About" column.
|
|
|
|
**Navigation is deliberately left alone.** `Navigation.tsx` renders a bottom tab
|
|
bar on mobile that already carries four items (Search, Compare, Rankings,
|
|
Admissions). A fifth tab makes each one cramped at 320px, and About and Blog are
|
|
both lower-intent than any of the four. Both live in the footer; About
|
|
additionally gets a byline link from every post, which is where a reader who
|
|
cares actually asks the question. Revisit only if analytics show people hunting
|
|
for it.
|
|
|
|
## Testing
|
|
|
|
Unit (Jest):
|
|
|
|
- Post rendering, including a post with no hero image and one with no excerpt.
|
|
- Slug generation and collision handling.
|
|
- JSON-LD shape for `BlogPosting` and `Person`.
|
|
- The `next.config.mjs` conversion preserves the staging `X-Robots-Tag` rule —
|
|
this guards the riskiest mechanical change in the plan.
|
|
|
|
E2E (Playwright, `e2e/`, required by CLAUDE.md for user-facing change):
|
|
|
|
- `/about` renders, shows the author name and photo, and is reachable from the
|
|
footer and nav.
|
|
- `/blog` lists at least one post; clicking through reaches the post.
|
|
- A post page renders title, date, body and byline.
|
|
- `/admin` responds with `noindex` and does not leak a stack trace when
|
|
unauthenticated.
|
|
|
|
Note the known constraint: new journeys cannot be proven in PR checks, because
|
|
the staging E2E gate runs post-merge.
|
|
|
|
## Risks
|
|
|
|
| Risk | Mitigation |
|
|
|---|---|
|
|
| `next.config.mjs` conversion silently drops the staging noindex header, making staging a crawlable duplicate | Unit test asserting the header rule; verify on staging before promotion |
|
|
| Payload API route collides with the FastAPI `/api` proxy | Remap to `/cms-api`; add to the proxy's exclusion list |
|
|
| Media volume mounts root-owned; all uploads fail with EACCES | `mkdir`+`chown` in the Dockerfile before the mount; test an upload on staging |
|
|
| Build fails or bakes empty content because CI has no DB | No build-time DB access; ISR only |
|
|
| A pipeline `--drop` destroys blog content | Separate `payload` schema; verify `--drop` blast radius before building |
|
|
| Media volume not backed up; images unrecoverable | Add `payload_media` to the backup routine |
|
|
| Payload auth CVE exposed publicly | Prompt patching; proxy restriction available as a fallback |
|
|
| Blog launches empty or goes stale | Ship with one post; cadence is explicitly "a few times a year", so no cadence is promised anywhere on the page — no dates implying a schedule |
|
|
|
|
## Sequence
|
|
|
|
Each step is independently reviewable and mergeable.
|
|
|
|
1. **Payload foundation** — install, `next.config.mjs` conversion, `payload`
|
|
schema, `/cms-api` remap, `users` collection, `/admin` hardening, compose and
|
|
Dockerfile changes. No public-facing change yet. Verify on staging that the
|
|
site is unchanged and `/admin` works.
|
|
2. **`/about`** — coded page, photo, `Person`/`Organization` JSON-LD, footer and
|
|
nav links, e2e journey. Independently valuable and does not depend on the
|
|
blog.
|
|
3. **Blog** — `posts` and `media` collections, `/blog` index and post pages, ISR
|
|
plus revalidation hooks, RSS, sitemap, structured data, e2e journeys.
|
|
4. **First post** — written in the admin panel, published through the normal
|
|
flow, proving the whole path end to end.
|
|
|
|
Step 1 carries all the infrastructure risk and none of the visible benefit, so
|
|
it should be verified on staging carefully before step 2 starts.
|
|
|
|
## Dependencies on Tudor
|
|
|
|
- **A photograph.** Blocks step 2. Nothing else in the plan is blocked by it.
|
|
- **The first post's subject.** Blocks step 4 only. Suggested: what school
|
|
performance data cannot tell you — it demonstrates judgement, is genuinely
|
|
useful, and is the kind of thing an anonymous or machine-written site will not
|
|
publish.
|
|
- ~~Confirmation that `scripts/migrate_csv_to_db.py --drop` is schema-scoped.~~
|
|
**Resolved 2026-09-02** — verified in `backend/migration.py`; see the
|
|
Database section. No action needed.
|