[E00-S03-T06] App does not report ready before migrations complete #395

Merged
kpcto merged 3 commits from feature/181 into main 2026-08-30 02:09:51 +00:00
Member

What changed

Implements [E00-S03-T06] App does not report ready before migrations complete (#181): apps/server now runs the startup migrations through the driver boundary (@personal-blog/database-postgres — the MigrationRunner over the migration ledger, E00-S03-T03) and gates its readiness on the run — the app does not report ready before migrations complete.

  • apps/server/src/index.ts — readiness gate (the acceptance criteria): GET /health is the readiness probe. A migrationsComplete flag starts false and flips to true only inside the startup migration run's success handler (runner.run().then(...)). While the run is in flight, /health answers HTTP 503 with {"status":"not ready"}; once the run finishes it answers HTTP 200 with {"status":"ok"}. When no DATABASE_URL is configured (the local non-container developer path, E00-S01-T06) there is no migration run to wait for, so the app reports ready immediately — keeping the E00-S02-T03 health endpoint and pnpm --filter @personal-blog/server start working. On a failed run the app logs the failure (the runner already throws the serializable MigrationFailedError from E00-S03-T05 — failure diagnostics are out of scope here) and stays not-ready, so a deployment with failed migrations is surfaced by the readiness probe instead of crash-looping.
  • apps/server/package.json + pnpm-lock.yaml — first cross-package dependency: the server now depends on @personal-blog/database-postgres (workspace:*); the lockfile importer was regenerated by pnpm. pg/Kysely stay isolated to database-postgres (E00-S03-T02) — the server imports only the boundary re-exports.
  • apps/server/package.json scripts — clean-clone typecheck/build: the server's build/typecheck scripts first build their workspace dependency (pnpm --filter @personal-blog/database-postgres build && tsc -p tsconfig.json [--noEmit]), so pnpm run build/pnpm run typecheck work from a clean clone (no committed dist/).
  • apps/server/Dockerfile: the build stage now copies the database-postgres manifest (so the frozen in-image install matches the lockfile importers) and its source (so the server's self-building script compiles it in-image); the runtime stage ships packages/database-postgres/dist + package.json so the server's @personal-blog/database-postgres import resolves through the copied workspace node_modules links.
  • tests/app-readiness.test.mjs — new suite locking in both acceptance criteria: static assertions on the committed entrypoint (driver-boundary import, 503 not-ready payload, if (migrationsComplete) gate in the /health route, flip-after-run() ordering, no-DATABASE_URL ready-immediately path), each backed by mutation probes proving non-vacuity; two deterministic behavioral probes boot the committed server over real HTTP (no DATABASE_URL → 200 {"status":"ok"} immediately; unreachable DATABASE_URL → the app stays up but answers 503 {"status":"not ready"}, never 200); the docker-gated real-stack probe executes the issue's test plan — "start with pending migrations and confirm readiness waits" — against the committed compose db: it holds an ACCESS EXCLUSIVE lock on the migration ledger so the app's startup migration run is genuinely pending, asserts /health stays 503 not-ready while the run is blocked, then releases the lock and asserts /health flips to 200 {"status":"ok"} once the run completes (and the app logs the completed run).
  • .gitea/workflows/ci.yml — new app-readiness job runs node --test tests/app-readiness.test.mjs on every PR (installs the frozen workspace + builds database-postgres, since the probes boot the committed server from the host). Additive only (no existing job modified, same action majors, no secrets: context, no untrusted interpolation).
  • Existing tests/docs: tests/compose-config.test.mjs — the real-stack health probe now treats a transient 503 (the app is up but not ready while migrations run) as "still starting — retry" instead of fail-fast, since the health endpoint is now the readiness probe; tests/health-endpoint.test.mjs — the smoke boot strips DATABASE_URL so it deterministically exercises the committed no-migration readiness path (200 immediately); docs/development/non-container.md and the runner.ts module header document the gate.

Explicitly out of scope per the brief, not touched: failure diagnostic (E00-S03-T05 — the app merely consumes the runner's existing error) and migration ledger (E00-S03-T03 — the app uses the ledger through the runner, no ledger changes).

Criterion → test table

Acceptance criterion Test (fails without the committed state)
app does not report ready before migrations complete tests/app-readiness.test.mjs — "the /health route reports not-ready while migrations are pending and ready only after they finish" (if (migrationsComplete) gates 200 vs sendJson(res, 503, NOT_READY_PAYLOAD) in the /health route; let migrationsComplete = false; starts not-ready); mutation probes "removing the 503 not-ready branch …", "answering 200 in the not-ready state …", "replacing the gate with an unconditional healthy answer …", "removing the not-ready payload …", "flipping the readiness flag before the migration run …", "dropping the migration run …"; deterministic probe "the app does not report ready while the migration run cannot complete (unreachable database)" — the committed server boots with a dead DATABASE_URL and answers 503 {"status":"not ready"} on every poll, never 200, and stays up; real-stack probe — with the migration run blocked behind a held ACCESS EXCLUSIVE lock on the ledger, /health answers 503 {"status":"not ready"} on every poll (the issue's test plan: "start with pending migrations and confirm readiness waits")
readiness is reported only after migrations finish tests/app-readiness.test.mjs — "the readiness flag flips to true only inside the startup migration run success handler" (in the DATABASE_URL path, migrationsComplete = true sits inside runner.run().then(...) at a source index after the runner.run() call; the runner is built as new MigrationRunner(pool, MIGRATIONS, new MigrationLedger(pool))); real-stack probe — after the blocked run completes (lock released), /health answers 200 {"status":"ok"} and the app logs startup migration run complete (applied 0, skipped 0), so the flip is attributable to a completed run; deterministic probe "without DATABASE_URL the app reports ready immediately" — the no-migration path answers 200 {"status":"ok"} (keeps the E00-S02-T03 health endpoint working)
the gate runs migrations through the driver boundary (pg/Kysely stay isolated to database-postgres) tests/app-readiness.test.mjs — "the server runs migrations through the driver boundary, never importing pg/Kysely directly" (import { Pool } / import { MigrationLedger, MigrationRunner } / import type { Migration } from @personal-blog/database-postgres; doesNotMatch from 'pg' / from 'kysely') + mutation probe "importing pg directly instead of the driver boundary …"; reinforced by the existing tests/database-postgres-imports.test.mjs isolation scan over the committed tree
the criteria are verified behaviorally — the real-stack probe is the authoritative check, static assertions are backed by mutation probes tests/app-readiness.test.mjs — the docker-gated real-stack probe runs the committed server against the committed compose db service and asserts both acceptance criteria behaviorally (503 while the run is pending → 200 after it completes, with the completion log and the intact migration ledger); the deterministic probes (no Docker) cover the no-DATABASE_URL and unreachable-DATABASE_URL contracts on every CI Node; every static assertion has a mutation probe (9 probes: 8 source mutations + placeholder module) proving it fails on a violation
the readiness criterion gates merges via the app-readiness job in .gitea/workflows/ci.yml tests/app-readiness.test.mjs — "the app-readiness criterion is enforced in CI" (root test glob covers the suite and .gitea/workflows/ci.yml runs it); .gitea/workflows/ci.yml — app-readiness job runs node --test tests/app-readiness.test.mjs on every PR (additive, matching the security-reviewed #390/#391/#392/#393/#394 precedent)

Test plan executed

  • node --test tests/app-readiness.test.mjs → 18 tests, 17 pass / 0 fail / 1 skip on Node 22.23.2 (the docker-gated real-stack probe skips — no Docker daemon in this sandbox; it runs in CI on Node 24 with Docker). The two deterministic probes ran for real: no DATABASE_URL → HTTP 200 {"status":"ok"}; unreachable DATABASE_URL → persistent HTTP 503 {"status":"not ready"}, app stays up.
  • Full suite (node --test "tests/**/*.test.mjs"): 198 tests — 174 pass / 9 fail / 15 skip on Node 22.23.2; the 9 failures are pre-existing environment artifacts identical to the base commit (this sandbox has Node 22 — the workspace engines gate requires Node ≥ 24): tests/frozen-install.test.mjs ×5 and tests/root-commands.test.mjs ×3 fail on ERR_PNPM_UNSUPPORTED_ENGINE (verified: the committed lockfile passes pnpm install --frozen-lockfile under the engine override, and root pnpm run build/pnpm run typecheck pass from a clean state under the override), tests/node-engine.test.mjs ×1 asserts the runtime is Node 24.x. CI runs Node 24 where these pass.
  • Frozen install + lockfile: pnpm install --frozen-lockfile succeeds against the regenerated pnpm-lock.yaml (apps/server importer gains @personal-blog/database-postgres: workspace:* → link:../../packages/database-postgres).
  • Typecheck/build from clean state: with all dist/ removed, pnpm run build (4/4 packages) and pnpm run typecheck (4/4 packages) exit 0 — the server's self-building scripts produce database-postgres/dist before compiling the server.
  • Related suites re-run locally, all green: health-endpoint (7), compose-config (17 pass / 4 docker-skip), build-targets (11 pass / 2 docker-skip), secrets-not-embedded (19 pass / 1 docker-skip), database-postgres-imports (8), architecture-import (10), no-core-extension-imports (3), workspace-layout (8), workspace-config (6), typescript-pin (3), strict-tsconfig (4), database-postgres-ledger (12 pass / 1 docker-skip), database-postgres-lock (16 pass / 2 skip), database-postgres-diagnostic (16 pass / 2 skip).

Risks / notes

  • The real-stack probe blocks the app's migration run with a held ACCESS EXCLUSIVE lock on the migration ledger (from a background psql session it kills to release) — deterministic, no sleeps: the test waits for the lock to be held (pg_locks poll) before booting the app and polls /health until it answers.
  • The compose-config real-stack health probe now retries on transient 503 (the app is up but not ready while migrations run) and fail-fasts only on a definitive non-2xx answer — required for the readiness-gated app and harmless otherwise.
  • Rollback note from the issue: revert the readiness gating logic — drop the migrationsComplete gate from apps/server/src/index.ts and the app's dependency on @personal-blog/database-postgres (reverting the gate alone restores the T03/T05 behavior; the migration wiring and the suite are additive).

Refs #181

## What changed Implements [E00-S03-T06] App does not report ready before migrations complete (#181): `apps/server` now runs the startup migrations through the driver boundary (`@personal-blog/database-postgres` — the `MigrationRunner` over the migration ledger, E00-S03-T03) and gates its readiness on the run — the app does **not** report ready before migrations complete. - **`apps/server/src/index.ts` — readiness gate (the acceptance criteria)**: `GET /health` is the readiness probe. A `migrationsComplete` flag starts `false` and flips to `true` only inside the startup migration run's success handler (`runner.run().then(...)`). While the run is in flight, `/health` answers HTTP 503 with `{"status":"not ready"}`; once the run finishes it answers HTTP 200 with `{"status":"ok"}`. When no `DATABASE_URL` is configured (the local non-container developer path, E00-S01-T06) there is no migration run to wait for, so the app reports ready immediately — keeping the E00-S02-T03 health endpoint and `pnpm --filter @personal-blog/server start` working. On a failed run the app logs the failure (the runner already throws the serializable `MigrationFailedError` from E00-S03-T05 — failure diagnostics are out of scope here) and stays not-ready, so a deployment with failed migrations is surfaced by the readiness probe instead of crash-looping. - **`apps/server/package.json` + `pnpm-lock.yaml` — first cross-package dependency**: the server now depends on `@personal-blog/database-postgres` (`workspace:*`); the lockfile importer was regenerated by pnpm. `pg`/Kysely stay isolated to `database-postgres` (E00-S03-T02) — the server imports only the boundary re-exports. - **`apps/server/package.json` scripts — clean-clone typecheck/build**: the server's `build`/`typecheck` scripts first build their workspace dependency (`pnpm --filter @personal-blog/database-postgres build && tsc -p tsconfig.json [--noEmit]`), so `pnpm run build`/`pnpm run typecheck` work from a clean clone (no committed `dist/`). - **`apps/server/Dockerfile`**: the build stage now copies the `database-postgres` manifest (so the frozen in-image install matches the lockfile importers) and its source (so the server's self-building script compiles it in-image); the runtime stage ships `packages/database-postgres/dist` + `package.json` so the server's `@personal-blog/database-postgres` import resolves through the copied workspace node_modules links. - **`tests/app-readiness.test.mjs` — new suite locking in both acceptance criteria**: static assertions on the committed entrypoint (driver-boundary import, 503 not-ready payload, `if (migrationsComplete)` gate in the `/health` route, flip-after-`run()` ordering, no-DATABASE_URL ready-immediately path), each backed by mutation probes proving non-vacuity; two **deterministic behavioral probes** boot the committed server over real HTTP (no DATABASE_URL → 200 `{"status":"ok"}` immediately; unreachable DATABASE_URL → the app stays up but answers 503 `{"status":"not ready"}`, never 200); the docker-gated **real-stack probe** executes the issue's test plan — "start with pending migrations and confirm readiness waits" — against the committed compose `db`: it holds an ACCESS EXCLUSIVE lock on the migration ledger so the app's startup migration run is genuinely pending, asserts `/health` stays 503 not-ready while the run is blocked, then releases the lock and asserts `/health` flips to 200 `{"status":"ok"}` once the run completes (and the app logs the completed run). - **`.gitea/workflows/ci.yml` — new `app-readiness` job** runs `node --test tests/app-readiness.test.mjs` on every PR (installs the frozen workspace + builds `database-postgres`, since the probes boot the committed server from the host). **Additive only** (no existing job modified, same action majors, no `secrets:` context, no untrusted interpolation). - **Existing tests/docs**: `tests/compose-config.test.mjs` — the real-stack health probe now treats a transient 503 (the app is up but not ready while migrations run) as "still starting — retry" instead of fail-fast, since the health endpoint is now the readiness probe; `tests/health-endpoint.test.mjs` — the smoke boot strips `DATABASE_URL` so it deterministically exercises the committed no-migration readiness path (200 immediately); `docs/development/non-container.md` and the `runner.ts` module header document the gate. Explicitly out of scope per the brief, **not touched**: failure diagnostic (E00-S03-T05 — the app merely consumes the runner's existing error) and migration ledger (E00-S03-T03 — the app uses the ledger through the runner, no ledger changes). ## Criterion → test table | Acceptance criterion | Test (fails without the committed state) | | --- | --- | | app does not report ready before migrations complete | `tests/app-readiness.test.mjs` — **"the /health route reports not-ready while migrations are pending and ready only after they finish"** (`if (migrationsComplete)` gates 200 vs `sendJson(res, 503, NOT_READY_PAYLOAD)` in the `/health` route; `let migrationsComplete = false;` starts not-ready); mutation probes **"removing the 503 not-ready branch …"**, **"answering 200 in the not-ready state …"**, **"replacing the gate with an unconditional healthy answer …"**, **"removing the not-ready payload …"**, **"flipping the readiness flag before the migration run …"**, **"dropping the migration run …"**; deterministic probe **"the app does not report ready while the migration run cannot complete (unreachable database)"** — the committed server boots with a dead `DATABASE_URL` and answers 503 `{"status":"not ready"}` on every poll, never 200, and stays up; real-stack probe — with the migration run blocked behind a held ACCESS EXCLUSIVE lock on the ledger, `/health` answers 503 `{"status":"not ready"}` on every poll (the issue's test plan: "start with pending migrations and confirm readiness waits") | | readiness is reported only after migrations finish | `tests/app-readiness.test.mjs` — **"the readiness flag flips to true only inside the startup migration run success handler"** (in the DATABASE_URL path, `migrationsComplete = true` sits inside `runner.run().then(...)` at a source index after the `runner.run()` call; the runner is built as `new MigrationRunner(pool, MIGRATIONS, new MigrationLedger(pool))`); real-stack probe — after the blocked run completes (lock released), `/health` answers 200 `{"status":"ok"}` and the app logs `startup migration run complete (applied 0, skipped 0)`, so the flip is attributable to a completed run; deterministic probe **"without DATABASE_URL the app reports ready immediately"** — the no-migration path answers 200 `{"status":"ok"}` (keeps the E00-S02-T03 health endpoint working) | | the gate runs migrations through the driver boundary (pg/Kysely stay isolated to database-postgres) | `tests/app-readiness.test.mjs` — **"the server runs migrations through the driver boundary, never importing pg/Kysely directly"** (`import { Pool }` / `import { MigrationLedger, MigrationRunner }` / `import type { Migration }` from `@personal-blog/database-postgres`; `doesNotMatch` `from 'pg'` / `from 'kysely'`) + mutation probe **"importing pg directly instead of the driver boundary …"**; reinforced by the existing `tests/database-postgres-imports.test.mjs` isolation scan over the committed tree | | the criteria are verified behaviorally — the real-stack probe is the authoritative check, static assertions are backed by mutation probes | `tests/app-readiness.test.mjs` — the docker-gated **real-stack probe** runs the committed server against the committed compose `db` service and asserts both acceptance criteria behaviorally (503 while the run is pending → 200 after it completes, with the completion log and the intact migration ledger); the **deterministic probes** (no Docker) cover the no-DATABASE_URL and unreachable-DATABASE_URL contracts on every CI Node; every static assertion has a **mutation probe** (9 probes: 8 source mutations + placeholder module) proving it fails on a violation | | the readiness criterion gates merges via the `app-readiness` job in `.gitea/workflows/ci.yml` | `tests/app-readiness.test.mjs` — **"the app-readiness criterion is enforced in CI"** (root test glob covers the suite and `.gitea/workflows/ci.yml` runs it); `.gitea/workflows/ci.yml` — **`app-readiness` job** runs `node --test tests/app-readiness.test.mjs` on every PR (additive, matching the security-reviewed #390/#391/#392/#393/#394 precedent) | ## Test plan executed - `node --test tests/app-readiness.test.mjs` → **18 tests, 17 pass / 0 fail / 1 skip** on Node 22.23.2 (the docker-gated real-stack probe skips — no Docker daemon in this sandbox; it runs in CI on Node 24 with Docker). The two deterministic probes ran for real: no `DATABASE_URL` → HTTP 200 `{"status":"ok"}`; unreachable `DATABASE_URL` → persistent HTTP 503 `{"status":"not ready"}`, app stays up. - **Full suite** (`node --test "tests/**/*.test.mjs"`): **198 tests — 174 pass / 9 fail / 15 skip on Node 22.23.2**; the 9 failures are pre-existing environment artifacts identical to the base commit (this sandbox has Node 22 — the workspace engines gate requires Node ≥ 24): `tests/frozen-install.test.mjs` ×5 and `tests/root-commands.test.mjs` ×3 fail on `ERR_PNPM_UNSUPPORTED_ENGINE` (verified: the committed lockfile passes `pnpm install --frozen-lockfile` under the engine override, and root `pnpm run build`/`pnpm run typecheck` pass from a clean state under the override), `tests/node-engine.test.mjs` ×1 asserts the runtime is Node 24.x. CI runs Node 24 where these pass. - **Frozen install + lockfile**: `pnpm install --frozen-lockfile` succeeds against the regenerated `pnpm-lock.yaml` (`apps/server` importer gains `@personal-blog/database-postgres: workspace:* → link:../../packages/database-postgres`). - **Typecheck/build from clean state**: with all `dist/` removed, `pnpm run build` (4/4 packages) and `pnpm run typecheck` (4/4 packages) exit 0 — the server's self-building scripts produce `database-postgres/dist` before compiling the server. - Related suites re-run locally, all green: `health-endpoint` (7), `compose-config` (17 pass / 4 docker-skip), `build-targets` (11 pass / 2 docker-skip), `secrets-not-embedded` (19 pass / 1 docker-skip), `database-postgres-imports` (8), `architecture-import` (10), `no-core-extension-imports` (3), `workspace-layout` (8), `workspace-config` (6), `typescript-pin` (3), `strict-tsconfig` (4), `database-postgres-ledger` (12 pass / 1 docker-skip), `database-postgres-lock` (16 pass / 2 skip), `database-postgres-diagnostic` (16 pass / 2 skip). ## Risks / notes - The real-stack probe blocks the app's migration run with a held ACCESS EXCLUSIVE lock on the migration ledger (from a background psql session it kills to release) — deterministic, no sleeps: the test waits for the lock to be held (`pg_locks` poll) before booting the app and polls `/health` until it answers. - The compose-config real-stack health probe now retries on transient 503 (the app is up but not ready while migrations run) and fail-fasts only on a definitive non-2xx answer — required for the readiness-gated app and harmless otherwise. - Rollback note from the issue: revert the readiness gating logic — drop the `migrationsComplete` gate from `apps/server/src/index.ts` and the app's dependency on `@personal-blog/database-postgres` (reverting the gate alone restores the T03/T05 behavior; the migration wiring and the suite are additive). Refs #181
bot-implementer added 3 commits 2026-08-30 01:56:50 +00:00
docs: document the readiness gate in the non-container guide (E00-S03-T06)
CI / Frozen lockfile install (pull_request) Successful in 49s
CI / Secrets not embedded (E00-S02-T08) (pull_request) Successful in 23s
CI / Database-postgres import isolation (E00-S03-T02) (pull_request) Successful in 26s
CI / Migration advisory lock (E00-S03-T04) (pull_request) Successful in 40s
CI / Migration failure diagnostic (E00-S03-T05) (pull_request) Successful in 56s
CI / Migration ledger (E00-S03-T03) (pull_request) Successful in 41s
CI / App readiness after migrations (E00-S03-T06) (pull_request) Successful in 55s
CI / Compose config (E00-S03-T01) (pull_request) Successful in 25s
5b0bded4b4
kpcto merged commit 6a65d4c789 into main 2026-08-30 02:09:51 +00:00
kpcto deleted branch feature/181 2026-08-30 02:09:51 +00:00
Sign in to join this conversation.