Runbooks
Runbooks are the reviewed operator path for repeated maintenance, recovery, and high-signal proof work. They should stay short enough to follow under pressure, but explicit about approval gates, mutation scope, and evidence.
The CLI command surface is documented in the Wavemap CLI Command Reference. This page explains when and how the deployed-dev operations are used. Use the Deploy Dev and Smoke command reference sections for command ownership, wrapper paths, compatibility scripts, flags, and default mutation posture.
Deployed-Dev Boundary
Section titled “Deployed-Dev Boundary”These runbooks apply only to the shared deployed dev environment:
| Boundary | Current Value |
|---|---|
| Public app URL | https://dev.wavemap.app |
| Pulumi stack | aws-dev |
| Runtime region | us-east-2 |
| Runtime target | Sleepable EC2 host running Docker Compose |
| Runtime config store | AWS SSM Parameter Store under /wavemap/dev/runtime/* |
| Last-good receipt | /opt/wavemap/deployment/last-successful-runtime-release.json |
Durable examples use the public CLI form:
pnpm wavemap -- <command> [flags...]Root scripts such as pnpm deploy:dev:runtime and pnpm smoke:dev remain compatibility aliases. Use them when a local
note or workflow already does, but prefer public wavemap routes in new runbooks.
Examples that need live Pulumi stack outputs use this placeholder path:
/tmp/wavemap-aws-dev-outputs.jsonCapture or supply that file through the approved deploy workflow, cloud-plan job, or local operator process before running commands that need live cloud facts. Do not paste secret values into the file or into runbook notes.
Shared Operator Gates
Section titled “Shared Operator Gates”Before running a live command, confirm:
- The target is deployed
dev, not local Docker, staging, or production. - The selected checkout/ref is the one intended for deployment or verification.
- Any live command has been dry-run first when the command supports a dry-run posture.
- The command’s mutation gate is explicit, usually
--execute. - Secrets are not printed, copied into workflow summaries, or stored in GitHub when the runtime host should read them from SSM.
- The expected evidence is known before starting: SSM command ID, GitHub job summary, smoke result, discrepancy counts, or failure artifact.
Operator Store Quick Reference
Section titled “Operator Store Quick Reference”Use this table before deciding which credential, artifact, or workflow path to reach for. The durable store and access boundary is explained in Infrastructure Change Policy.
| Task | Normal Input | Actor | Gate |
|---|---|---|---|
| Manual application-only delivery | Exact merged, CI-green SHA plus current deployment contract and application role. | .github/workflows/deploy-application-dev.yml. | Changed-path admission must remain application-only; no migration, seed, reset, or Pulumi authority. |
| Reviewed deployed-dev operation | Selected merged ref and one deterministic operator profile. | .github/workflows/deploy-dev.yml. | Human-selected recipe; mutating, destructive, media, and lifecycle stages retain their explicit gates. |
| Manual docs publish | Selected reviewed ref plus current deployment contract and docs role. | .github/workflows/deploy-docs.yml. | Docs role reads the private contract store. No Pulumi token or application-runtime authority. |
| Infrastructure mutation | Pulumi backend, Pulumi config, and infra operator credentials. | Local infra operator, or future explicitly approved infra workflow. | Clean full-stack saved plan, exact plan/source digests, human approval, and post-apply evidence. |
| Manual infra-topology ingest | Tendril /wavemap backend selector, backend region, prefix-scoped OIDC role, passphrase. | .github/workflows/infra-topology-ingest.yml or local operator capture. | Read-only access through the exact protected infra-topology environment; no workload mutation authority. |
| Runtime secret value change | SSM Parameter Store path under /wavemap/dev/runtime/*. | Runtime config population command or approved operator process. | GitHub should carry references or bootstrap values, not decrypted runtime secrets. |
| Deployment contract publication | Reviewed Pulumi outputs projected into wavemap.deployment-contract v1. | Approved contract publisher. | Future writer permission still needs least-privilege review before receiving artifact-store write access. |
| Public topology figure publication | Reviewed sanitized topology projection and figure ledger decision. | Human docs edit after private capture/projection review. | Raw captures and private generated candidates stay outside public docs. |
Guarded Pulumi Operations And Recovery
Section titled “Guarded Pulumi Operations And Recovery”The active aws-dev state authority is the Tendril state foundation’s dedicated /wavemap namespace. The former
workload-account backend is retained recovery evidence, not a current migration target. Do not reconstruct historical
Pulumi Cloud or direct-import commands from old notes.
All ordinary previews and updates use bin/infra/pulumi-stack.sh through the infra/pulumi package scripts. The launcher
fixes the project, stack, Tendril backend, Tooling and workload identities, workload profile, region, passphrase file,
source state, and saved-plan boundary.
Before a cloud-aware preview:
- Use a clean checkout at the exact source revision being reviewed.
- Confirm the Tendril Tooling and Wavemap workload profiles resolve to their expected accounts.
- Store the stack passphrase at
.local/secrets/aws-dev-passphrasewith mode600. - Keep the saved plan under an ignored private path with mode
600and a restrictiveumask. - Confirm no other infrastructure operation or state recovery is active.
Create one full-stack, refresh-enabled saved plan:
umask 077export WAVEMAP_INFRA_PLAN_PATH=/absolute/private/path/to/aws-dev.planpnpm -C infra/pulumi run preview -- --refresh --save-plan "$WAVEMAP_INFRA_PLAN_PATH"shasum -a 256 "$WAVEMAP_INFRA_PLAN_PATH"git rev-parse HEADReview the plan as a graph transition. Record the exact plan SHA-256 and full source Git SHA. An update requires a separate authorization and applies only that reviewed plan:
export WAVEMAP_INFRA_APPROVED_PLAN_SHA256=<reviewed-plan-sha256>export WAVEMAP_INFRA_APPROVED_SOURCE_GIT_SHA=<reviewed-full-source-git-sha>pnpm -C infra/pulumi run up -- \ --approved-plan-sha256 "$WAVEMAP_INFRA_APPROVED_PLAN_SHA256" \ --approved-source-git-sha "$WAVEMAP_INFRA_APPROVED_SOURCE_GIT_SHA" \ --plan "$WAVEMAP_INFRA_PLAN_PATH" \ --skip-preview \ --yesThe launcher rejects targeted or dirty saved plans, mismatched plan/source digests, alternate stack/backend/config selection, inline passphrases, ambiguous Pulumi environment variables, and changed file-backed object inputs. Targeted previews remain diagnostic only and cannot create an update plan.
If state, history, or a pending operation appears inconsistent:
- Stop mutation and preserve the current backend, checkpoint, Pulumi history, and provider readback.
- Classify whether the discrepancy is an active operation, an abandoned pending operation, a checkpoint/history issue, or workload drift.
- Compare the Tendril
/wavemapauthority with the retained legacy recovery evidence without importing either one over the other. - Choose forward repair, checkpoint recovery, or provider reconciliation only after the exact evidence and blast radius are reviewed.
- Re-run a guarded refresh-enabled full-stack preview and independently read the checkpoint, history, pending-operation, and workload baselines afterward.
Never clear a lock, import a checkpoint, expose secrets, or switch to the legacy backend merely because a command timed
out or a Pulumi history entry says failed. Provider postconditions and checkpoint state must be reconciled first.
Source-reading path:
bin/infra/pulumi-stack.showns runtime guards and environment selection.bin/infra/audit-pulumi-saved-plan.mjsowns saved-plan identity and graph-shape checks.infra/pulumi/index.tsandinfra/pulumi/src/providers/awsown the Wavemap workload graph.infra/pulumi/src/providers/aws/__tests__/pulumiStackWrapper.test.tsowns the launcher regression contract.- Tendril
infra/foundation/src/state-backend.tsowns the shared backend resource and namespace layout.
GitHub Workflow Dispatch Recipe Selection
Section titled “GitHub Workflow Dispatch Recipe Selection”Use this runbook when manually dispatching .github/workflows/deploy-dev.yml for deployed dev.
The dispatch posture is recipe-only. Select one deterministic run_profile; the workflow does not expose raw stage
toggles or arbitrary stage composition.
Before dispatch:
- Confirm the selected
git_refis the branch or SHA intended for this proof. - Confirm whether the run should be non-mutating, app-runtime mutating, data-destructive, media-mutating, or lifecycle-disruptive.
- Confirm a heavier profile is worth its cost, runtime disruption, and artifact noise.
- Decide what evidence will make the run useful before starting it.
Profile selection:
| Profile | Use When | Boundary |
|---|---|---|
preflight | Checking repo-local deploy contracts through GitHub Actions. | No cloud authentication and no live mutation. |
preflight-docker | Checking the same contracts plus local Docker image builds. | No cloud authentication or push; slower than ordinary preflight. |
cloud-plan | Checking the deployment contract, GitHub OIDC, live SSM metadata, and deploy dry-runs. | Cloud-authenticated but non-mutating. |
deploy-endpoint | Running a reviewed migration-aware app/API operator deploy. | Builds/pushes images, deploys runtime, repairs pending migrations, and gates endpoint smoke. |
deploy-endpoint-recovery | Adding non-destructive wake and browser-routing recovery proof to an endpoint deploy. | No reset or media mutation. |
deploy-seeded-browser | Proving the seeded route and browser basics after app or data-shape changes. | Includes destructive database reset before seeded and browser smoke. |
deploy-media | Proving API, S3, public media URL, CloudFront media delivery, browser rendering, and drift counts. | Includes destructive reset, temporary media mutation, browser media proof, and DB/S3 report. |
deploy-lifecycle | Proving shutdown, cold-start page, wake, and browser reload behavior. | Includes destructive reset and deliberately stops the runtime host for recovery proof. |
deploy-full-validation | Deliberately combining endpoint recovery, seeded, media, and lifecycle proof. | Broadest and most disruptive recipe; includes reset, media mutation, and host stop. |
Keep deploy-media and deploy-lifecycle separate for normal use. Select deploy-full-validation only when the
combined proof is deliberate.
Expected evidence:
- Workflow run number or URL.
- Selected ref, resolved commit SHA, and selected deterministic profile.
- Resolved stage summary.
- SSM command IDs for runtime deploy, database status/migrate, reset, release-record, rollback, or discrepancy-report stages when they run.
- Smoke results, elapsed times, and artifact names for any browser, media, or lifecycle failures.
Use local commands for focused manual repair, dry-run planning, or when a workflow job’s summary points at a specific operator action.
Runtime And Data Lifecycle Quick Reference
Section titled “Runtime And Data Lifecycle Quick Reference”The deployed-dev runtime is cost-first and disposable, but different operations have different data consequences. Use Deployed Dev Lifecycle for the full lifecycle matrix, teardown gradations, and expected evidence. Use Data Durability And Recovery for the current disposable-data posture, backup learning-drill boundary, and future recovery gates. Use Media Workflow And Validation when choosing between media smoke, browser media smoke, and discrepancy reporting.
| Operation | Runtime Behavior | Database Behavior | Media Behavior | Operator Gate |
|---|---|---|---|---|
| Runtime deploy | Pulls selected backend/frontend images and restarts Compose on the current host. | Preserved unless the deploy also runs an explicit reset or migration path. | Unchanged. | runtime deploy --execute after images and runtime config are ready. |
| Host stop or automatic inactivity shutdown | Stops the EC2 instance; root EBS remains attached. | Preserved across stop/start. | Unchanged. | Shutdown Lambda or approved lifecycle proof. |
| Runtime rollback | Redeploys the last-good app image pair through the runtime deploy document. | Not rolled back. | Not rolled back. | runtime rollback --execute, followed by endpoint smoke. |
| Database status | Runs a read-only migration ledger check in the deployed API container. | Read-only. | Unchanged. | database status, before migration repair or deploy schema gates. |
| Database migrate | Runs pending Drizzle migrations in the deployed API container. | Schema mutation only; existing rows should be preserved. | Unchanged. | database migrate --execute, followed by status and endpoint smoke. |
| Database reset | Verifies and clears guarded DB schemas, then migrates and reseeds in the API container. | Destructive; existing rows are removed before migrations run. | Unchanged; objects can become application-orphaned. | database reset --execute, followed by seeded smoke. |
| Media discrepancy report | Runs a read-only DB/S3 comparison in the deployed API container. | Read-only. | Read-only. | media discrepancy-report --execute for live SSM execution. |
| EC2/runtime replacement | Replaces or destroys the host through infrastructure change. | At risk unless a separate backup, snapshot, or migration runbook is used. | Unchanged unless media infrastructure also changes. | Runtime host replacement after approval. |
| Media bucket replacement or destroy | Not a runtime-host action. | Rows may reference missing media after replacement/delete. | At risk. | Media bucket replacement after approval. |
Host stop is cost control. It is not database cleanup, media cleanup, backup, rollback, or deploy-state mutation.
Host Stop And Wake Recovery
Section titled “Host Stop And Wake Recovery”Use this runbook when the deployed-dev runtime host is stopped, may be stopped, or needs a deliberate stopped-host recovery proof.
Host stop is a shared-environment disruption. Prefer the deploy-lifecycle workflow profile when the goal is a normal
stopped-host proof, because the workflow captures shutdown, cold-start, browser, and smoke evidence in one place. Use the
manual checks below for focused recovery, control-plane debugging, or when an automatic inactivity stop has already
occurred.
Before deliberate host stop:
- Confirm the target is
https://dev.wavemap.app, Pulumi stackaws-dev, and runtime regionus-east-2. - Confirm no deploy, reset, media proof, or demo is in progress.
- Confirm whether endpoint wake recovery is enough, or whether the browser must return to its original destination.
- Confirm the seeded baseline is valid before choosing browser cold-start recovery.
- Confirm the stop is only cost-control or lifecycle proof. Do not combine it with database reset, media cleanup, rollback, backup, replacement, or stack teardown without a separate operator decision.
Deliberate host stop is owned by the shutdown Lambda. It has no public app route. This is a live cloud mutation and
needs explicit operator approval at run time. If a local operator invokes it outside the deploy-lifecycle workflow,
derive the shutdown function name from the approved deployment contract’s resource prefix or another reviewed infra
output, then use an approved AWS operator identity:
aws lambda invoke \ --region us-east-2 \ --function-name "<shutdownFunctionName>" \ /tmp/wavemap-dev-shutdown-response.jsonAfter stopping the host, confirm the app URL serves the cold-start page rather than a raw CloudFront, connection, or app error. The cold-start page should appear at the original app URL while the runtime host is stopped or warming up.
For endpoint wake recovery, run:
pnpm wavemap -- smoke dev --wakeThis calls the same-origin /__wake path, waits for frontend readiness, and then replays endpoint smoke. Use it when the
host may simply be asleep, when validating the wake Lambda and dynamic origin refresh, or when the seeded browser
baseline is not known to be valid.
For browser destination recovery from an intentionally stopped host, run:
pnpm wavemap -- smoke dev cold-start-browserThis starts at the seeded Sorsari artist route, expects the cold-start page at that original URL, lets the page call
/__wake, waits for readiness, verifies the browser reloads into the original destination, and then replays endpoint and
seeded checks.
If the host was stopped automatically by the inactivity monitor, treat recovery the same way:
- Use
pnpm wavemap -- smoke dev --wakefor cheap endpoint recovery. - Use
pnpm wavemap -- smoke dev cold-start-browseronly when the seeded baseline is still valid and browser recovery is the evidence you need. - Do not add a database reset only to make browser recovery easier; reset remains a separate destructive decision.
If runtime deploy starts while the host is stopped, the runtime deploy wrapper should wake the host and retry SSM while the instance comes back online. Use the Runtime Deploy runbook and capture resume telemetry from the deploy output.
Expected evidence:
- Shutdown Lambda response or workflow stage summary when a deliberate stop was performed.
- Confirmation that the stopped app URL served the cold-start page.
- Wake smoke result or cold-start browser smoke result.
- Resume telemetry if runtime deploy performed the wake.
- Playwright artifacts, shutdown response, cold-start precheck HTML, or workflow summary links when recovery fails.
Failure triage:
| Symptom | First Place To Look |
|---|---|
| Cold-start page does not appear for the stopped app URL. | CloudFront custom error behavior, app-origin connection timeout settings, and static cold-start origin. |
| Wake call succeeds but readiness times out. | Wake Lambda logs, EC2 instance state, dynamic app-origin.dev.wavemap.app DNS update, and app startup logs. |
| Runtime deploy reports early SSM target errors. | Resume telemetry; retryable SSM registration delay is expected while a stopped host comes back. |
| Endpoint wake passes but browser recovery fails. | Cold-start page client reload behavior, original destination routing, and Playwright artifacts. |
Runtime Host Replacement
Section titled “Runtime Host Replacement”Use this runbook when a Pulumi preview proposes replacing or destroying the deployed-dev EC2 runtime host, or when an operator intentionally chooses to recreate the runtime host for AMI, bootstrap, host-size, VPC, subnet, security-group, or instance-profile work.
Runtime host replacement is not host stop, runtime deploy, runtime rollback, or database reset. It is an infrastructure change that can discard the containerized Postgres data path and the host-side last-good runtime release receipt.
Before approval:
- Confirm the target is deployed
dev, Pulumi stackaws-dev, and regionus-east-2. - Confirm no deploy, reset, media proof, or demo is in progress.
- Run Pulumi preview and identify every create, replacement, delete, IAM change, DNS change, SSM document change, and runtime output change.
- Confirm whether the preview changes only runtime-host resources or also touches media storage, CloudFront, DNS, IAM, or SSM runtime configuration.
- Decide the database outcome before applying: accept data loss and reseed, preserve through a separate backup/restore drill, or abort the replacement.
- Decide whether the host-side last-good release receipt matters. Replacement can remove it, so plan to record a fresh receipt after a known-good deploy.
- Confirm runtime app secrets and config remain owned by SSM Parameter Store and do not need to be copied from the old host.
Preview the infrastructure change:
pnpm -C infra/pulumi run previewApply only after explicit approval for the replacement and data outcome:
pnpm -C infra/pulumi run upAfter apply, capture fresh stack outputs:
pnpm -C infra/pulumi exec pulumi stack output --stack aws-dev --json > /tmp/wavemap-aws-dev-outputs.jsonValidate the captured target and runtime contracts:
pnpm wavemap -- deploy dev cloud-target --pulumi-outputs-json /tmp/wavemap-aws-dev-outputs.jsonpnpm wavemap -- deploy dev runtime-config live --pulumi-outputs-json /tmp/wavemap-aws-dev-outputs.jsonpnpm wavemap -- deploy dev runtime-env --pulumi-outputs-json /tmp/wavemap-aws-dev-outputs.jsonDeploy an intended image pair to the new host:
pnpm wavemap -- deploy dev bundle --pulumi-outputs-json /tmp/wavemap-aws-dev-outputs.jsonpnpm wavemap -- deploy dev runtime deploy --pulumi-outputs-json /tmp/wavemap-aws-dev-outputs.json --execute --github-outputpnpm wavemap -- smoke devIf the chosen database outcome was “disposable reset”, reseed and prove the seeded baseline:
pnpm wavemap -- deploy dev database reset --pulumi-outputs-json /tmp/wavemap-aws-dev-outputs.json --execute --github-outputpnpm wavemap -- smoke dev seededAfter endpoint smoke passes, record a fresh last-good receipt:
pnpm wavemap -- deploy dev runtime record-release --pulumi-outputs-json /tmp/wavemap-aws-dev-outputs.json --execute --github-outputExpected evidence:
- Pulumi preview summary naming the runtime-host replacement and any dependent resource changes.
- Explicit approval note covering the database outcome.
- Refreshed deployment contract after apply.
- Cloud-target, runtime-config, and runtime-env readiness results.
- Runtime deploy SSM command ID against the new instance.
- Endpoint smoke result, and seeded smoke result if reset was selected.
- Fresh last-good release receipt after the new host is proven.
Failure triage:
| Symptom | First Place To Look |
|---|---|
| Runtime deploy cannot target the instance. | Refreshed deployment contract, EC2 instance state, SSM managed-instance registration, and instance profile. |
| SSM command runs but Docker or Compose setup fails. | EC2 bootstrap user data, Docker service status, Compose plugin install, ECR login, and /opt/wavemap paths. |
| App URL does not reach the new host. | Dynamic app-origin.dev.wavemap.app record, CloudFront app origin settings, security group ingress, and port 80. |
| App starts but seeded routes are missing. | Database outcome decision, Postgres data path, migration/reset output, and seeded smoke artifacts. |
| Rollback cannot find a receipt. | Expected after host replacement unless a fresh receipt has been recorded on the new host. |
Media Bucket Replacement Or Destroy
Section titled “Media Bucket Replacement Or Destroy”Use this runbook when a Pulumi preview proposes replacing, destroying, or recreating the deployed-dev media S3 bucket or
its delivery path. This includes changes that alter the bucket physical name, CloudFront media origin, bucket policy,
origin access control, runtime media outputs, or runtime MEDIA_S3_BUCKET_NAME value.
The deployed-dev media bucket is intentionally disposable and currently uses forced cleanup at the infrastructure level. That makes replacement possible, not routine. Bucket replacement can delete objects and can leave database rows pointing at media that no longer exists.
Before approval:
- Confirm the target is deployed
dev, Pulumi stackaws-dev, and regionus-east-2. - Confirm no media smoke, upload test, database reset, or demo is in progress.
- Run Pulumi preview and identify every media bucket, bucket policy, CloudFront, IAM, runtime config, and output change.
- Decide the object-data outcome before applying: accept deletion, preserve through a separate object-copy plan, or abort.
- Decide the database outcome before applying: keep rows as-is, run destructive reset after replacement, or plan a separate row migration for locator fields.
- Remember that S3 media rows store
storageLocation/thumbnailStorageLocation; copying objects to a new bucket may still require a DB locator migration if old rows should remain deletable and reconcilable. - Confirm GitHub Actions will not decrypt runtime media secrets; the runtime host reads media config through SSM-rendered env files.
If current DB/S3 drift matters, capture a read-only baseline before replacement:
pnpm wavemap -- deploy dev media discrepancy-report --pulumi-outputs-json /tmp/wavemap-aws-dev-outputs.json --execute --github-outputPreview the infrastructure change:
pnpm -C infra/pulumi run previewApply only after explicit approval for object deletion/preservation and database handling:
pnpm -C infra/pulumi run upAfter apply, capture fresh stack outputs:
pnpm -C infra/pulumi exec pulumi stack output --stack aws-dev --json > /tmp/wavemap-aws-dev-outputs.jsonRefresh runtime media configuration and re-render the host env through runtime deploy:
pnpm wavemap -- deploy dev cloud-target --pulumi-outputs-json /tmp/wavemap-aws-dev-outputs.jsonpnpm wavemap -- deploy dev runtime-config plan --pulumi-outputs-json /tmp/wavemap-aws-dev-outputs.jsonpnpm wavemap -- deploy dev runtime-config populate --pulumi-outputs-json /tmp/wavemap-aws-dev-outputs.json --executepnpm wavemap -- deploy dev runtime-config live --pulumi-outputs-json /tmp/wavemap-aws-dev-outputs.jsonpnpm wavemap -- deploy dev runtime deploy --pulumi-outputs-json /tmp/wavemap-aws-dev-outputs.json --execute --github-outputThen prove app and media behavior:
pnpm wavemap -- smoke devpnpm wavemap -- smoke dev media --pulumi-outputs-json /tmp/wavemap-aws-dev-outputs.json --executepnpm wavemap -- smoke dev browser-mediapnpm wavemap -- deploy dev media discrepancy-report --pulumi-outputs-json /tmp/wavemap-aws-dev-outputs.json --execute --github-outputIf the chosen database outcome was “disposable reset”, run the reset before media/browser-media proof:
pnpm wavemap -- deploy dev database reset --pulumi-outputs-json /tmp/wavemap-aws-dev-outputs.json --execute --github-outputpnpm wavemap -- smoke dev seededExpected evidence:
- Pulumi preview summary naming the bucket replacement/delete and dependent CloudFront/IAM/runtime-config changes.
- Explicit approval note covering object handling and database handling.
- Optional pre-replacement discrepancy report when preserving or interpreting existing media matters.
- Fresh Pulumi outputs capture after apply.
- Runtime-config plan/populate/live evidence showing the new media bucket value is ready without printing secrets.
- Runtime deploy SSM command ID proving the host env was re-rendered.
- Endpoint smoke, media smoke, browser media smoke, and post-replacement discrepancy report.
Failure triage:
| Symptom | First Place To Look |
|---|---|
| Upload fails after replacement. | Runtime MEDIA_S3_BUCKET_NAME, runtime host role media policy, bucket existence, and bucket public-access block. |
| Upload succeeds but public media 403s. | CloudFront media origin, bucket policy for origin access control, path-prefix strip function, and object key. |
| Existing media rows render missing images. | Object handling decision, DB locator fields, copied object keys, and discrepancy report output. |
| Delete or cleanup targets the old bucket. | Stored storageLocation / thumbnailStorageLocation values and any planned locator migration. |
| Browser-media smoke fails but API smoke passes. | CloudFront /media/* routing, image URL returned by the API, browser test artifacts, and cache behavior. |
Runtime Config Readiness
Section titled “Runtime Config Readiness”Use this runbook when a deploy is blocked by missing runtime configuration, a runtime parameter changed, or an operator needs to verify SSM parameter readiness before runtime deploy.
This is the operational procedure for runtime config readiness. Source ownership lives in Configuration And Secrets, and the deployed-dev environment shape lives in Deployed Dev Environment. There is no separate runtime-config operations page until the operator surface grows beyond this runbook.
Runtime config source ownership:
- Pulumi outputs expose parameter names and value kinds, not plaintext secrets.
- SSM Parameter Store owns deployed app runtime config and runtime secrets.
- GitHub Actions should not decrypt or log
SecureStringruntime values. - Frontend
NEXT_PUBLIC_*values are browser build inputs, not runtime SSM parameters.
This readiness check owns the application parameter groups declared by runtime-config-contract.ts: backend, database, frontend, i18n and media. It rejects missing required parameters, wrong types and unknown names inside those groups. Other namespaces beneath the shared runtime prefix remain outside this check, including the public reader and origin credentials owned by Public API Runtime. A passing application readiness result does not establish those credentials or authorize API activation.
For the broader source-ownership and secret-handling convention, see Configuration And Secrets.
Plan the required runtime configuration without cloud calls:
pnpm wavemap -- deploy dev runtime-config planpnpm wavemap -- deploy dev runtime-config plan --pulumi-outputs-json /tmp/wavemap-aws-dev-outputs.jsonPlan SSM population for non-secret String parameters:
pnpm wavemap -- deploy dev runtime-config populate --pulumi-outputs-json /tmp/wavemap-aws-dev-outputs.jsonWrite only the selected non-secret String parameters after the operator accepts the plan:
pnpm wavemap -- deploy dev runtime-config populate --pulumi-outputs-json /tmp/wavemap-aws-dev-outputs.json --executeCheck live SSM metadata without reading parameter values:
pnpm wavemap -- deploy dev runtime-config live --pulumi-outputs-json /tmp/wavemap-aws-dev-outputs.jsonVerify the host env rendering contract without contacting the runtime host:
pnpm wavemap -- deploy dev runtime-env --pulumi-outputs-json /tmp/wavemap-aws-dev-outputs.jsonStop if readiness is blocked. Fix the missing name/type/source problem first, then rerun the readiness check. Secret
creation or rotation is a separate operator action; do not use runtime-config populate to create SecureString
parameters.
Expected evidence:
- Runtime config plan groups, required/optional counts, and redacted secret placeholders.
- Live metadata readiness showing required parameters present with expected
String/SecureStringtypes. - No decrypted secret values in logs, local files, workflow summaries, or artifacts.
Runtime Deploy
Section titled “Runtime Deploy”Use this runbook to deploy an already selected app-runtime image pair to deployed dev. The reviewed operator workflow
uses deploy-endpoint; local runtime deploy is mainly for focused repair, replaying a deploy after images already exist,
or proving the runtime handoff directly. Manual application-only delivery uses the separate database-free digest lane.
Preconditions:
- Runtime config readiness is green.
- Backend and frontend images for the selected
sha-<full-git-sha>tag exist in ECR, or the image build/push step will run before runtime deploy. - The deployment facts file points at the intended account, stack, region, runtime instance, and SSM document. In application delivery, this is the sanitized deployment contract read from the private artifact store.
- Database reset, media proof, and lifecycle proof are selected only if they are part of the intended recipe.
Plan the runtime bundle:
pnpm wavemap -- deploy dev bundle --pulumi-outputs-json /tmp/wavemap-aws-dev-outputs.jsonDry-run the runtime deploy command:
pnpm wavemap -- deploy dev runtime deploy --pulumi-outputs-json /tmp/wavemap-aws-dev-outputs.jsonExecute the runtime deploy:
pnpm wavemap -- deploy dev runtime deploy --pulumi-outputs-json /tmp/wavemap-aws-dev-outputs.json --execute --github-outputThen run endpoint smoke:
pnpm wavemap -- smoke devThe live runtime deploy sends the Pulumi-modeled SSM document to the runtime host. If the host is stopped, the wrapper uses the same-origin wake path and retries SSM while the instance comes back online. The deploy document renders env files from SSM Parameter Store, logs into ECR on-host, pulls the selected images, and starts Docker Compose detached.
Expected evidence:
- Runtime deploy SSM command ID.
- Resume telemetry when the host needed to wake.
- Successful command completion or a failure log that identifies the SSM send, wait, ECR pull, env render, or Compose stage.
- Endpoint smoke success after deploy.
Last-Good Runtime Release Receipt
Section titled “Last-Good Runtime Release Receipt”The last-good receipt is the rollback input for deployed dev. It should be written only after endpoint smoke succeeds
for the current image pair.
The reviewed operator workflow records this host-side receipt after successful endpoint smoke. The database-free application lane records its separate versioned digest receipt in S3. Use the local command only when manually repairing or reestablishing the operator rollback point after a known-good deploy.
Dry-run the receipt write:
pnpm wavemap -- deploy dev runtime record-release --pulumi-outputs-json /tmp/wavemap-aws-dev-outputs.jsonWrite the receipt after smoke has passed:
pnpm wavemap -- deploy dev runtime record-release --pulumi-outputs-json /tmp/wavemap-aws-dev-outputs.json --execute --github-outputThe receipt records the source git SHA, immutable image tag, backend image URI, frontend image URI, deployment version, selected ref, and workflow run metadata. It lives on the runtime host at:
/opt/wavemap/deployment/last-successful-runtime-release.jsonDo not write this receipt for a deploy that has not passed endpoint smoke. Doing so would teach rollback to restore an unproven image pair.
Expected evidence:
- Receipt-write SSM command ID.
- Receipt path.
- Source SHA, deployment version, and backend/frontend image tags matching the smoke-passing deploy.
Runtime Rollback
Section titled “Runtime Rollback”Runtime rollback is an app-container rollback only. It reads the host-side last-good receipt and redeploys that recorded backend/frontend image pair through the existing runtime deploy document.
Rollback does not roll back:
- Database schema or rows.
- SSM runtime parameters.
- Media objects.
- Pulumi infrastructure.
- Destructive reset outcomes.
Use this operator rollback when a reviewed runtime deploy or endpoint smoke failure leaves deployed dev on a bad tagged
image pair and a host-side receipt exists. The separate application-only lane performs its automatic rollback from the
previous versioned digest receipt instead.
Dry-run the rollback target:
pnpm wavemap -- deploy dev runtime rollback --pulumi-outputs-json /tmp/wavemap-aws-dev-outputs.jsonExecute rollback:
pnpm wavemap -- deploy dev runtime rollback --pulumi-outputs-json /tmp/wavemap-aws-dev-outputs.json --execute --github-outputThen rerun endpoint smoke:
pnpm wavemap -- smoke devIf rollback smoke fails, keep the failed release visible in the workflow summary and choose the next operator action explicitly: forward fix, manual deploy of a known SHA, data reset, or infrastructure investigation.
Expected evidence:
- Rollback SSM command ID.
- Restored source SHA, image tag, deployment version, and receipt path.
- Endpoint smoke result after rollback.
- Clear note if rollback was skipped because no last-good receipt existed.
Endpoint Diagnostics
Section titled “Endpoint Diagnostics”Use this runbook before repair work when deployed dev is reachable but an application route appears slow, stuck, or
generically broken. The command captures public HTTP evidence only; it does not use AWS credentials, SSM, Docker,
Pulumi, or database access.
Capture baseline evidence:
pnpm wavemap -- deploy dev diagnostics endpointsInclude seeded artist details when seeded data should exist:
pnpm wavemap -- deploy dev diagnostics endpoints --seeded-artistUse a stable evidence directory when the transcript needs to be attached to a workflow note or investigation:
pnpm wavemap -- deploy dev diagnostics endpoints --evidence-dir /tmp/wavemap-endpoint-diagnosticsExpected evidence:
- Evidence directory path.
- Per-endpoint headers, body, curl timing metadata, and curl stderr files.
- HTTP status, response content type, CloudFront cache header, and backend
x-request-idwhen present. - Short JSON body previews for failing or diagnostic API routes.
Interpretation posture:
/en/pingpassing means CloudFront can reach the frontend container./api/v1/healthpassing means the frontend rewrite can reach the backend process./api/v1/readypassing means the backend can reach Postgres for a shallowSELECT 1check.- Application route failures after readiness passes usually point at app logic, schema compatibility, data shape, or downstream service behavior rather than a stopped host.
Database Migration Status And Repair
Section titled “Database Migration Status And Repair”The reviewed deploy-endpoint operator recipe runs this same posture after runtime deploy and before endpoint smoke:
read-only status first, typed gate decision second, migration-only repair only when the live DB is behind, then status
again before smoke. Manual use is still useful for incident repair or educational diagnosis.
Use this runbook when deployed dev is reachable but DB-backed routes fail with schema errors, or before a migration-only
repair. This path is narrower than reset: it compares migration state and can run pending migrations without seeding or
deleting rows.
Capture live stack outputs first when needed:
pnpm -C infra/pulumi exec pulumi stack output --stack aws-dev --json > /tmp/wavemap-aws-dev-outputs.jsonRun the read-only migration status check:
pnpm wavemap -- deploy dev database statusStatus reads the deployed API image’s Drizzle journal and the live database’s drizzle.__drizzle_migrations ledger, then
emits status markers such as:
WAVEMAP_DATABASE_STATUS_STATE=behindWAVEMAP_DATABASE_STATUS_EXPECTED_COUNT=21WAVEMAP_DATABASE_STATUS_APPLIED_COUNT=13WAVEMAP_DATABASE_STATUS_PENDING_COUNT=8Interpretation posture:
up-to-date: The live database matches the deployed API image’s migration journal.behind: The live database is missing migrations present in the deployed API image. Prefer migration-only repair.ahead: The database has migration ledger rows newer than this deployed image. Stop and avoid applying older code.hash-mismatchordrift: Migration history may have been rewritten or skipped. Stop and inspect before repair.ledger-missingorempty: Treat as a fresh or damaged DB state; do not assume migration-only repair is safe without reviewing the data posture.
CD converts those states into a gate decision before endpoint smoke:
pass: Status isup-to-date; endpoint smoke may run.migrate: Status isbehindand migration repair is enabled; rundatabase migrate, recheck status, then smoke.block: Status is missing,behindwithout migration permission, or any drift/ahead/history-risk state; stop for manual review.
Dry-run the migration-only command:
pnpm wavemap -- deploy dev database migrate --pulumi-outputs-json /tmp/wavemap-aws-dev-outputs.jsonApply pending migrations only:
pnpm wavemap -- deploy dev database migrate --pulumi-outputs-json /tmp/wavemap-aws-dev-outputs.json --execute --github-outputThen verify:
pnpm wavemap -- deploy dev database statuspnpm wavemap -- deploy dev diagnostics endpointspnpm wavemap -- smoke devExpected evidence:
- Before/after database status output.
- SSM command ID for any
database migrate --executerun. - Endpoint diagnostics or smoke success proving the original failing route now works.
Database Reset
Section titled “Database Reset”Use this runbook only for deployed dev. The database is disposable, but reset is still an explicit destructive action.
Decision points before reset:
- Confirm the target is
https://dev.wavemap.app, Pulumi stackaws-dev, and deployment environmentdev. - Confirm losing non-seed database rows is acceptable.
- Confirm existing S3 media objects may outlive reset because reset does not delete the media bucket.
- Decide whether this is a normal disposable reset or a backup/restore learning drill.
- Prefer
deploy-seeded-browser,deploy-media, ordeploy-lifecycle; usedeploy-full-validationonly when the combined destructive proof is deliberate.
Dry-run the reset:
pnpm wavemap -- deploy dev database reset --pulumi-outputs-json /tmp/wavemap-aws-dev-outputs.jsonExecute the reset:
pnpm wavemap -- deploy dev database reset --pulumi-outputs-json /tmp/wavemap-aws-dev-outputs.json --execute --github-outputThen prove the canonical seed baseline:
pnpm wavemap -- smoke dev seededThe reset planner accepts only the AWS dev deployment contract and renders a dry-run by default. During execution, the
backend reset command requires the deployed-dev context, NODE_ENV=production, and the exact runtime database identity
(db:5432, database/user wavemap_dev). It clears the application and Drizzle migration schemas before running
migrations, base seed, and deterministic dev-data seed inside the deployed API container. Seeded smoke is the minimum
proof that the canonical seeded route and API state exist again.
This sequencing is intentionally different from migration-only repair: database migrate preserves retained rows, while
database reset rebuilds the disposable database from empty schemas. The reset uses the existing SSM command path and
does not require additional cloud API or IAM permissions.
The canonical first seeded app route is:
/en/artist/ef839db3-ae41-4af9-9078-a8d211089962Known reset warning posture:
- Historical live reset proof reported event-series import warnings.
- Treat those warnings as seed-data cleanup work, not as a failed reset, when migrations, base seed, deterministic dev-data seed, and seeded smoke all pass.
- If the warning shape changes, treat it as fresh evidence.
Expected evidence:
- Database reset SSM command ID.
- Reset command success or failure log.
- Seeded smoke success.
- Follow-up media discrepancy report when DB/S3 divergence matters.
Media Cleanup
Section titled “Media Cleanup”Media deletion can commit while provider cleanup remains pending. Use the shared backend command to inspect and retry originals and thumbnails for any owner family. Run from the environment configured for the intended database and each pending object’s stored provider. The command deletes storage objects and pending entries; it never deletes an entity.
Start with a read-only preview:
pnpm wavemap -- maintenance media cleanuppnpm wavemap -- maintenance media cleanup --owner-type events --owner-public-id WmEvent0001 --limit 25Owner types are artists, events, venues, event-series, event-series-edition and users. An owner public ID requires
its owner type; the same ID can belong to different entity families. Omit both filters to inspect all pending work.
The default limit is 100 objects, with an allowed range of 1–1000. Reports include before/after pending counts, selected,
deleted, failed and skipped counts, plus cleanup IDs, owner types/public IDs and retry diagnostics. Storage keys,
locations and URLs are omitted.
After checking the target and preview, execute the reviewed scope:
pnpm wavemap -- maintenance media cleanup --owner-type events --owner-public-id WmEvent0001 --limit 25 --executeThe compatibility script is pnpm maintenance:media:cleanup; both forms delegate to
pnpm -C apps/wavemap-back-end media:cleanup. Omission of --execute never contacts storage for deletion.
The command attempts one bounded snapshot, prioritizing unattempted entries and then the oldest attempt. Successful
objects are acknowledged by removing their pending entry; failures remain and produce a nonzero exit after the report.
Concurrent attempts skip locked entries. Remaining work can reflect the limit or another attempt, so inspect the report
before another pass.
Resolve provider availability or credentials and retry failed work. For a partial or invalid locator, use the cleanup ID
to inspect the corresponding media_cleanup row through restricted database access. Verify the intended object against
provider inventory before repairing its reference. Do not discard unresolved entries merely because they are old.
Provider success followed by a database acknowledgement failure is safe to retry; adapters accept already-missing objects.
There is no scheduled worker or automatic backlog purge. The operator owns follow-through until the intended scope’s
pendingAfter is zero. Capture the target environment, reviewed preview, execution result and remaining work. These rows
are pending recovery state, not a durable entity audit.
Upload compensation records failed objects when its independent database transaction succeeds. If recovery cannot be recorded, logs identify the owner and direct operators to discrepancy inspection. Legacy registration compensation runs inside the registration transaction and does not promise durable recovery. Read-only discrepancy reporting compares canonical media rows to provider objects; it does not subtract pending-cleanup entries. Compare both inventories before classifying an object as an unexplained orphan. Repair remains a separately reviewed action.
Event Media Cleanup
Section titled “Event Media Cleanup”The existing Event route and JSON report shape remain compatible. They now cover pending Event media work from both whole-Event and per-media operations through the shared cleanup service:
pnpm wavemap -- maintenance event media-cleanup --event-public-id WmEvent0001 --limit 25pnpm wavemap -- maintenance event media-cleanup --event-public-id WmEvent0001 --limit 25 --executepnpm maintenance:event:media-cleanup still delegates to pnpm -C apps/wavemap-back-end event:cleanup-media.
The alias scopes owner type to events and retains eventPublicId in its report entries. Existing pending Event work is
preserved by migration 0052_shared_media_cleanup. No live provider cleanup is part of implementation verification.
Venue Mutation Receipt Cleanup
Section titled “Venue Mutation Receipt Cleanup”Use this on-demand maintenance command to bound retry-receipt growth after the 30-day Venue mutation replay window. These rows protect idempotent create, update, and Event-relationship writes; they are operational safety state rather than a human-action audit trail. Private Venue address-transition audits are deliberately retained and are never targeted by this command.
Run the public maintenance preview from an environment configured for the exact database you intend to inspect:
pnpm wavemap -- maintenance venue mutation-receipts cleanupThe command is dry-run by default. Its report includes the exclusive cutoff, eligible rows before and after the run, oldest remaining receipt, maximum rows per run, and whether another bounded pass is needed. Defaults retain 30 days and inspect or delete at most ten batches of 1,000 rows. Override the bounds only when the review requires it:
pnpm wavemap -- maintenance venue mutation-receipts cleanup --retention-days 30 --batch-size 1000 --max-batches 10After confirming the database identity and preview, execute the reviewed deletion:
pnpm wavemap -- maintenance venue mutation-receipts cleanup --executeThe compatibility form is pnpm maintenance:venue:mutation-receipts. The public route delegates directly to the
backend-owned TypeScript entrypoint; the CLI does not duplicate retention or database behavior.
Each deletion batch uses deterministic oldest-first selection and its own short transaction. If
hasMoreEligibleReceipts is true, review the report and run another bounded pass rather than silently widening one run.
A future scheduler should invoke this same application-owned cleanup seam; it should not duplicate the retention query in
infrastructure code.
Expected evidence:
- Target environment and database identity.
- Dry-run cutoff, eligible count, and configured bounds.
- Execute report with rows deleted and batches executed.
- Remaining eligible count and whether another reviewed pass is required.
Media Discrepancy Report
Section titled “Media Discrepancy Report”Use this runbook after destructive reset, media smoke, manual media testing, or any investigation where database rows and S3 objects may have drifted.
Use Media Workflow And Validation for deciding whether a media change needs this report or a different proof lane.
The report is read-only. It should surface:
- Artist or Event media rows whose canonical or thumbnail locator points at an S3 object that no longer exists.
- S3 media objects under the deployed-dev media prefix that are no longer referenced by either entity-owned collection.
- Counts and sampled identifiers/keys without printing app secrets or signed credentials.
Dry-run the report:
pnpm wavemap -- deploy dev media discrepancy-report --pulumi-outputs-json /tmp/wavemap-aws-dev-outputs.jsonExecute the read-only report on the runtime host:
pnpm wavemap -- deploy dev media discrepancy-report --pulumi-outputs-json /tmp/wavemap-aws-dev-outputs.json --execute --github-outputThe live command sends an SSM command that runs the backend discrepancy report inside the deployed API container, so DB access and S3 listing use the rendered runtime environment. Artist and Event collection adapters contribute provider-neutral locators to the same evidence boundary. The command does not delete rows or objects.
Non-zero discrepancy counts are telemetry, not automatic cleanup and not automatically a deploy failure. Cleanup remains a separate approval-gated action that should print the exact DB rows or object keys it will touch before mutation.
Expected evidence:
- Media discrepancy SSM command ID.
- Report scope.
- Entity scopes inspected, currently Artist and Event media.
- Inspected row/object counts.
- Total discrepancy count and discrepancy-kind summary.
- Any sampled rows or object keys needed for follow-up.
Evidence And Failure Triage
Section titled “Evidence And Failure Triage”Use workflow summaries first. Deployed-dev workflow summaries and job summaries should identify the stable facts an operator needs without turning the docs site into a run archive:
- Selected deterministic profile.
- Selected ref, resolved commit SHA, and deployment version or image tag.
- Resolved stages.
- Runtime deploy, reset, rollback, release-record, and discrepancy SSM command IDs when those jobs run.
- Smoke result and elapsed time.
- First job log or artifact to inspect on failure.
Browser-style jobs upload Playwright artifacts on failure. Database reset and media smoke lanes upload failure-only logs when selected. Keep raw logs, command outputs, and live identifiers in private workflow artifacts unless they are intentionally sanitized for docs.
Use this routing rule when deciding where evidence belongs:
| Evidence Kind | Durable Home |
|---|---|
| Stable procedure, expected evidence, mutation boundary, failure triage, or redaction rule. | Curated docs under apps/wavemap-docs/src/content/docs. |
| Unsettled experiment, proof interpretation, temporary timing adjustment, or implementation log. | Working notes under apps/wavemap-docs/working-notes. |
| Change-specific proof that a PR, deploy, or workflow run behaved correctly. | Pull request description, workflow summary, or private workflow/job summary. |
| Raw command output, Lambda payload, browser trace, screenshot, downloaded artifact, or log. | Private artifacts with appropriate retention, unless intentionally sanitized before publication. |
| Run identifier, timestamp, commit SHA, SSM command ID, instance ID, public IP, or proof URL. | Private evidence, PR/workflow context, or working note while it is actively useful for follow-up. |
A live proof graduates to curated docs only after it has been reduced to the reusable lesson. For example, document that cold-start browser proof should capture the shutdown response, cold-start precheck HTML, Playwright failure artifacts, and final smoke result. Do not publish the specific workflow run number, EC2 instance ID, SSM command UUID, public IP, or timing measurement unless that identifier itself is part of a reviewed operator decision.
When evidence changes an operator path, update the runbook. When evidence only proves a single run, keep it with the run.
Publishing Policy
Section titled “Publishing Policy”Publish only reviewed operator paths, expected evidence rubrics, and sanitized examples that teach a stable pattern.
Do not publish:
- Raw logs, command output, Lambda responses, browser traces, screenshots, or downloaded artifacts.
- One-off proof narratives that do not change future operator behavior.
- Live cloud identifiers such as instance IDs, public IPs, SSM command UUIDs, or temporary proof URLs.
- Workflow run numbers, commit SHAs, and timestamps unless they are intentionally part of a historical decision record.
- Temporary proof acceleration details such as shortened inactivity windows.
Working notes may keep this material while a decision is still moving. Once the decision settles, promote the procedure, evidence expectation, or redaction rule into curated docs and leave the noisy proof trail in PRs, workflow summaries, or private artifacts.