[E00-S03-T06] App does not report ready before migrations complete #395
No Reviewers
Labels
Clear labels
agent/analyst-drafted
agent/analyst-drafted
needs/human-decision
needs/human-decision
needs/security-review
needs/security-review
tier/t0
tier/t1
tier/t2
tier/t3
kind
bug
kind
bug
kind
epic
kind
epic
kind
initiative
EPPP programme initiative
kind
story
kind
story
kind
task
EPPP engineering card/task decomposed from a story
kind
toil
kind
toil
loop
1
loop
1
loop
2
loop
2
loop
3
loop
3
risk
agent-full
risk
agent-full
risk
human-gated
risk
human-gated
risk
human-only
risk
human-only
size
l
size
l
size
m
size
m
size
s
size
s
status
blocked
status
blocked
status
done
Workflow: Done
status
in-progress
status
in-progress
status
proposed
status
proposed
status
ready
status
ready
status
review
status
review
stream
checkout
stream
checkout
stream
onboarding
stream
onboarding
stream
platform
stream
platform
trivial — implementer only, auto-merge
standard — implementer + reviewer + tester
complex — security if triggered, human merge
critical — full chain + security, human merge
No labels
Milestone
No items
No Milestone
Projects
Clear projects
No projects
No Assignees
Notifications
Due Date
No due date set.
Dependencies
No dependencies set.
Reference: Fabrika/PersonalBlog#395
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
What changed
Implements [E00-S03-T06] App does not report ready before migrations complete (#181):
apps/servernow runs the startup migrations through the driver boundary (@personal-blog/database-postgres— theMigrationRunnerover the migration ledger, E00-S03-T03) and gates its readiness on the run — the app does not report ready before migrations complete.apps/server/src/index.ts— readiness gate (the acceptance criteria):GET /healthis the readiness probe. AmigrationsCompleteflag startsfalseand flips totrueonly inside the startup migration run's success handler (runner.run().then(...)). While the run is in flight,/healthanswers HTTP 503 with{"status":"not ready"}; once the run finishes it answers HTTP 200 with{"status":"ok"}. When noDATABASE_URLis configured (the local non-container developer path, E00-S01-T06) there is no migration run to wait for, so the app reports ready immediately — keeping the E00-S02-T03 health endpoint andpnpm --filter @personal-blog/server startworking. On a failed run the app logs the failure (the runner already throws the serializableMigrationFailedErrorfrom E00-S03-T05 — failure diagnostics are out of scope here) and stays not-ready, so a deployment with failed migrations is surfaced by the readiness probe instead of crash-looping.apps/server/package.json+pnpm-lock.yaml— first cross-package dependency: the server now depends on@personal-blog/database-postgres(workspace:*); the lockfile importer was regenerated by pnpm.pg/Kysely stay isolated todatabase-postgres(E00-S03-T02) — the server imports only the boundary re-exports.apps/server/package.jsonscripts — clean-clone typecheck/build: the server'sbuild/typecheckscripts first build their workspace dependency (pnpm --filter @personal-blog/database-postgres build && tsc -p tsconfig.json [--noEmit]), sopnpm run build/pnpm run typecheckwork from a clean clone (no committeddist/).apps/server/Dockerfile: the build stage now copies thedatabase-postgresmanifest (so the frozen in-image install matches the lockfile importers) and its source (so the server's self-building script compiles it in-image); the runtime stage shipspackages/database-postgres/dist+package.jsonso the server's@personal-blog/database-postgresimport resolves through the copied workspace node_modules links.tests/app-readiness.test.mjs— new suite locking in both acceptance criteria: static assertions on the committed entrypoint (driver-boundary import, 503 not-ready payload,if (migrationsComplete)gate in the/healthroute, flip-after-run()ordering, no-DATABASE_URL ready-immediately path), each backed by mutation probes proving non-vacuity; two deterministic behavioral probes boot the committed server over real HTTP (no DATABASE_URL → 200{"status":"ok"}immediately; unreachable DATABASE_URL → the app stays up but answers 503{"status":"not ready"}, never 200); the docker-gated real-stack probe executes the issue's test plan — "start with pending migrations and confirm readiness waits" — against the committed composedb: it holds an ACCESS EXCLUSIVE lock on the migration ledger so the app's startup migration run is genuinely pending, asserts/healthstays 503 not-ready while the run is blocked, then releases the lock and asserts/healthflips to 200{"status":"ok"}once the run completes (and the app logs the completed run)..gitea/workflows/ci.yml— newapp-readinessjob runsnode --test tests/app-readiness.test.mjson every PR (installs the frozen workspace + buildsdatabase-postgres, since the probes boot the committed server from the host). Additive only (no existing job modified, same action majors, nosecrets:context, no untrusted interpolation).tests/compose-config.test.mjs— the real-stack health probe now treats a transient 503 (the app is up but not ready while migrations run) as "still starting — retry" instead of fail-fast, since the health endpoint is now the readiness probe;tests/health-endpoint.test.mjs— the smoke boot stripsDATABASE_URLso it deterministically exercises the committed no-migration readiness path (200 immediately);docs/development/non-container.mdand therunner.tsmodule header document the gate.Explicitly out of scope per the brief, not touched: failure diagnostic (E00-S03-T05 — the app merely consumes the runner's existing error) and migration ledger (E00-S03-T03 — the app uses the ledger through the runner, no ledger changes).
Criterion → test table
tests/app-readiness.test.mjs— "the /health route reports not-ready while migrations are pending and ready only after they finish" (if (migrationsComplete)gates 200 vssendJson(res, 503, NOT_READY_PAYLOAD)in the/healthroute;let migrationsComplete = false;starts not-ready); mutation probes "removing the 503 not-ready branch …", "answering 200 in the not-ready state …", "replacing the gate with an unconditional healthy answer …", "removing the not-ready payload …", "flipping the readiness flag before the migration run …", "dropping the migration run …"; deterministic probe "the app does not report ready while the migration run cannot complete (unreachable database)" — the committed server boots with a deadDATABASE_URLand answers 503{"status":"not ready"}on every poll, never 200, and stays up; real-stack probe — with the migration run blocked behind a held ACCESS EXCLUSIVE lock on the ledger,/healthanswers 503{"status":"not ready"}on every poll (the issue's test plan: "start with pending migrations and confirm readiness waits")tests/app-readiness.test.mjs— "the readiness flag flips to true only inside the startup migration run success handler" (in the DATABASE_URL path,migrationsComplete = truesits insiderunner.run().then(...)at a source index after therunner.run()call; the runner is built asnew MigrationRunner(pool, MIGRATIONS, new MigrationLedger(pool))); real-stack probe — after the blocked run completes (lock released),/healthanswers 200{"status":"ok"}and the app logsstartup migration run complete (applied 0, skipped 0), so the flip is attributable to a completed run; deterministic probe "without DATABASE_URL the app reports ready immediately" — the no-migration path answers 200{"status":"ok"}(keeps the E00-S02-T03 health endpoint working)tests/app-readiness.test.mjs— "the server runs migrations through the driver boundary, never importing pg/Kysely directly" (import { Pool }/import { MigrationLedger, MigrationRunner }/import type { Migration }from@personal-blog/database-postgres;doesNotMatchfrom 'pg'/from 'kysely') + mutation probe "importing pg directly instead of the driver boundary …"; reinforced by the existingtests/database-postgres-imports.test.mjsisolation scan over the committed treetests/app-readiness.test.mjs— the docker-gated real-stack probe runs the committed server against the committed composedbservice and asserts both acceptance criteria behaviorally (503 while the run is pending → 200 after it completes, with the completion log and the intact migration ledger); the deterministic probes (no Docker) cover the no-DATABASE_URL and unreachable-DATABASE_URL contracts on every CI Node; every static assertion has a mutation probe (9 probes: 8 source mutations + placeholder module) proving it fails on a violationapp-readinessjob in.gitea/workflows/ci.ymltests/app-readiness.test.mjs— "the app-readiness criterion is enforced in CI" (root test glob covers the suite and.gitea/workflows/ci.ymlruns it);.gitea/workflows/ci.yml—app-readinessjob runsnode --test tests/app-readiness.test.mjson every PR (additive, matching the security-reviewed #390/#391/#392/#393/#394 precedent)Test plan executed
node --test tests/app-readiness.test.mjs→ 18 tests, 17 pass / 0 fail / 1 skip on Node 22.23.2 (the docker-gated real-stack probe skips — no Docker daemon in this sandbox; it runs in CI on Node 24 with Docker). The two deterministic probes ran for real: noDATABASE_URL→ HTTP 200{"status":"ok"}; unreachableDATABASE_URL→ persistent HTTP 503{"status":"not ready"}, app stays up.node --test "tests/**/*.test.mjs"): 198 tests — 174 pass / 9 fail / 15 skip on Node 22.23.2; the 9 failures are pre-existing environment artifacts identical to the base commit (this sandbox has Node 22 — the workspace engines gate requires Node ≥ 24):tests/frozen-install.test.mjs×5 andtests/root-commands.test.mjs×3 fail onERR_PNPM_UNSUPPORTED_ENGINE(verified: the committed lockfile passespnpm install --frozen-lockfileunder the engine override, and rootpnpm run build/pnpm run typecheckpass from a clean state under the override),tests/node-engine.test.mjs×1 asserts the runtime is Node 24.x. CI runs Node 24 where these pass.pnpm install --frozen-lockfilesucceeds against the regeneratedpnpm-lock.yaml(apps/serverimporter gains@personal-blog/database-postgres: workspace:* → link:../../packages/database-postgres).dist/removed,pnpm run build(4/4 packages) andpnpm run typecheck(4/4 packages) exit 0 — the server's self-building scripts producedatabase-postgres/distbefore compiling the server.health-endpoint(7),compose-config(17 pass / 4 docker-skip),build-targets(11 pass / 2 docker-skip),secrets-not-embedded(19 pass / 1 docker-skip),database-postgres-imports(8),architecture-import(10),no-core-extension-imports(3),workspace-layout(8),workspace-config(6),typescript-pin(3),strict-tsconfig(4),database-postgres-ledger(12 pass / 1 docker-skip),database-postgres-lock(16 pass / 2 skip),database-postgres-diagnostic(16 pass / 2 skip).Risks / notes
pg_lockspoll) before booting the app and polls/healthuntil it answers.migrationsCompletegate fromapps/server/src/index.tsand the app's dependency on@personal-blog/database-postgres(reverting the gate alone restores the T03/T05 behavior; the migration wiring and the suite are additive).Refs #181