fix(blog): hide drafts at the access layer, and back the --drop claim
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m12s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 32s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m9s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 1m15s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 2m26s

Review findings on #140.

Drafts were reachable. Posts granted unconditional public read and the
_status filter lived only in the pages that query the collection — which
is a convenience, not a control. Payload's documentation is explicit:
"The `draft` argument alone does not restrict documents with _status:
'draft' from being returned by the API." A direct GET /cms-api/posts
would have handed every unpublished draft to any visitor. Read access
now returns a query constraint for anonymous callers, which is the
documented mechanism.

The --drop claim was asserted across four files while the spec still
listed it as an open question. Now verified rather than assumed:
run_full_migration drops exactly ["school_results", "schools"] by name,
there is no drop_all() or DROP SCHEMA anywhere in backend/, the only
other drop is schema-qualified to marts, and nothing sets search_path.
The guarantee is stronger than schema isolation alone — those two table
names do not exist in Payload — so the claim stands, but it now rests on
cited code. The spec records the evidence and closes the open item.

findPost is wrapped in React's cache(): Next calls generateMetadata and
the page separately for one request, so every post view ran the same
query against Postgres twice.

The bare .lede rule was dead — .prose p scores (0,1,1) and outranks it —
so only .prose .lede ever applied. Removed, with the specificity noted
so the surviving selector is not "simplified" back into a silent
regression.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017YmbBhr8s7GusjDE12hrZM
This commit is contained in:
TudorandClaude Opus 5 committed 2026-09-02 16:48:59 +01:00
1 parent 07d586d0ad
commit e25722d9ab
5 files changed
+67 -13

No files matched your search

@@ -152,8 +152,31 @@ reach `sc_database:5432` with no networking change. It needs a new
Schema isolation is not cosmetic. `public` currently holds the application Schema isolation is not cosmetic. `public` currently holds the application
tables and Airflow's metadata, and `scripts/migrate_csv_to_db.py --drop` exists tables and Airflow's metadata, and `scripts/migrate_csv_to_db.py --drop` exists
to drop and reimport. Blog content living in its own schema means no data to drop and reimport. Blog content living in its own schema means no data
pipeline operation can destroy it. **Before implementation, confirm that pipeline operation can destroy it.
`--drop` is schema-scoped and cannot reach `payload`.**
**Verified 2026-09-02** (this was an open question when the spec was written).
`--drop` calls `run_full_migration()` in `backend/migration.py`, which drops
exactly two tables by name:
```python
ks2_tables = ["school_results", "schools"]
for tname in ks2_tables:
if tname in existing:
Base.metadata.tables[tname].drop(bind=engine)
```
There is no `Base.metadata.drop_all()` anywhere in `backend/`, and no
`DROP SCHEMA`. The only other drop is `_apply_schema_drops()`, a single
schema-qualified `DROP TABLE IF EXISTS marts.fact_parent_view CASCADE`.
Nothing sets `search_path`, so the SQLAlchemy metadata resolves to `public`,
and `inspector.get_table_names()` does not even enumerate other schemas.
So the guarantee is stronger than schema isolation alone: `--drop` targets two
named tables that Payload does not have, and would not reach `posts`, `media`
or `users` even if they shared a schema. The `payload` schema remains the right
choice — it protects against a *future* broadening of that script rather than
today's behaviour — but the safety claim rests on verified code, not on
assumption.
Putting CMS tables in this instance is consistent with existing practice — Putting CMS tables in this instance is consistent with existing practice —
Airflow already stores its metadata there. Airflow already stores its metadata there.
@@ -338,4 +361,6 @@ it should be verified on staging carefully before step 2 starts.
performance data cannot tell you — it demonstrates judgement, is genuinely performance data cannot tell you — it demonstrates judgement, is genuinely
useful, and is the kind of thing an anonymous or machine-written site will not useful, and is the kind of thing an anonymous or machine-written site will not
publish. publish.
- **Confirmation** that `scripts/migrate_csv_to_db.py --drop` is schema-scoped. - ~~Confirmation that `scripts/migrate_csv_to_db.py --drop` is schema-scoped.~~
**Resolved 2026-09-02** — verified in `backend/migration.py`; see the
Database section. No action needed.
@@ -29,8 +29,14 @@ describe('posts collection', () => {
expect(slugField).toMatch(/index:\s*true/); expect(slugField).toMatch(/index:\s*true/);
}); });
it('is publicly readable', () => { it('hides drafts from anonymous readers at the access layer', () => {
expect(POSTS).toMatch(/access:\s*\{\s*read:\s*\(\)\s*=>\s*true/); // Payload's docs are explicit: "The `draft` argument alone does not
// restrict documents with _status: 'draft' from being returned by the
// API." The blog pages' where-clause is not enforcement — a direct GET
// /cms-api/posts would return unpublished drafts to anyone. Access
// control returning a query constraint is the only thing that stops it.
expect(POSTS).toMatch(/_status:\s*\{\s*equals:\s*'published'\s*\}/);
expect(POSTS).toMatch(/if\s*\(req\.user\)\s*return true/);
}); });
it('revalidates the post page when a post changes or is deleted', () => { it('revalidates the post page when a post changes or is deleted', () => {
@@ -54,12 +54,11 @@
} }
/* The opening paragraph carries the page. Larger, and in the primary ink /* The opening paragraph carries the page. Larger, and in the primary ink
rather than the secondary, so it reads as a voice rather than as body copy. */ rather than the secondary, so it reads as a voice rather than as body copy.
.lede {
font-size: 1.125rem;
color: var(--text-primary);
}
Must stay in the descendant form: `.prose p` scores (0,1,1) and would beat a
bare `.lede` at (0,1,0), so simplifying this selector silently reverts the
lede to ordinary body copy. */
.prose .lede { .prose .lede {
font-size: 1.125rem; font-size: 1.125rem;
color: var(--text-primary); color: var(--text-primary);
@@ -1,3 +1,4 @@
import { cache } from 'react';
import type { Metadata } from 'next'; import type { Metadata } from 'next';
import Link from 'next/link'; import Link from 'next/link';
import { notFound } from 'next/navigation'; import { notFound } from 'next/navigation';
@@ -50,7 +51,13 @@ const calloutConverters: JSXConvertersFunction = ({ defaultConverters }) => ({
}, },
}); });
async function findPost(slug: string) { /**
* Wrapped in React's cache() because Next calls generateMetadata and the page
* component separately for the same request — without it, every post view runs
* this query against Postgres twice. cache() dedupes within a single request
* only, so it never serves one visitor's request from another's.
*/
const findPost = cache(async (slug: string) => {
const payload = await getCachedPayload(); const payload = await getCachedPayload();
const { docs } = await payload.find({ const { docs } = await payload.find({
collection: 'posts', collection: 'posts',
@@ -59,7 +66,7 @@ async function findPost(slug: string) {
depth: 1, depth: 1,
}); });
return docs[0] ?? null; return docs[0] ?? null;
} });
function summarise(post: Record<string, unknown>) { function summarise(post: Record<string, unknown>) {
return { return {
+18 -1
View File
@@ -24,7 +24,24 @@ function revalidatePost(slug: string) {
export const Posts: CollectionConfig = { export const Posts: CollectionConfig = {
slug: 'posts', slug: 'posts',
access: { read: () => true }, access: {
/*
* Drafts must be hidden here, not in the pages that query this collection.
*
* From Payload's own documentation: "The `draft` argument alone does not
* restrict documents with `_status: 'draft'` from being returned by the
* API." The blog index and post page both filter on `_status`, but that
* is a convenience, not a control — a direct GET /cms-api/posts would
* hand every unpublished draft to any visitor.
*
* Returning a query constraint rather than a boolean is the documented
* mechanism: Payload merges it into every read for an anonymous caller.
*/
read: ({ req }) => {
if (req.user) return true;
return { _status: { equals: 'published' } };
},
},
admin: { admin: {
useAsTitle: 'title', useAsTitle: 'title',
defaultColumns: ['title', 'publishedAt', '_status'], defaultColumns: ['title', 'publishedAt', '_status'],