Public API Reader And Runtime
The public process receives one restricted database credential and public delivery configuration. Database provisioning, schema maintenance and shared privilege changes belong to a separately authorized operator. Public startup and ordinary application delivery never run those operations. This page documents reader maintenance, local proof and the modeled Development release and HTTPS boundary. Live secret population, deployment and traffic activation require separate operator approval.
Reader Contract
Section titled “Reader Contract”apps/wavemap-back-end/src/db/publicReader/contract.ts owns an explicit allowlist of canonical Drizzle column references, derives their physical SQL names, and includes two invoker functions: normalize_entity_name(text) and unaccent(text). The role receives database CONNECT, schema USAGE and explicit column SELECT; it receives no write, schema-create, temporary-table, sequence, private-table or unreviewed application-function access. Media storage locators needed to construct delivery URLs remain internal query inputs; authoring identity, original filename, MIME type, byte size and blur metadata are excluded. Lifecycle visibility and location-publication rules remain query/projection responsibilities; this is not row-level database isolation.
Schema property renames and removals fail typechecking at the selected references; physical SQL names follow the schema metadata. Adding a column does not grant access to it. A new public query dependency still requires an intentional privilege-policy change and real reader proof. Response DTO schemas cannot supply this policy because joins, filters, ordering and media URL resolution also consume internal database columns.
reconcile.ts plans and applies the contract in one transaction. It validates both the supplied target and connected identity, uses bounded lock/statement timeouts, and requires the guarded administrator. Reader names use wavemap_public_ followed by 1–40 lowercase letters, digits or underscores. Existing roles must carry the managed marker, have no elevated attributes, memberships, owned objects, outside-scope dependencies or session defaults, and have no active connections. Unexpected authority requires separate operator repair; the tool does not adopt arbitrary roles or run DROP OWNED.
Two exact target postures are supported: the repository’s disposable test identity on loopback port 55433 with NODE_ENV=test and context db-backed-test; or the existing container target db:5432/wavemap_dev, administrator wavemap_dev, NODE_ENV=production and context deployed-dev. Other targets, URL overrides and fragments are rejected before connecting. Changing this support boundary requires a reviewed implementation change.
Plan, Apply And Check
Section titled “Plan, Apply And Check”Use the public reader maintenance command, which routes through the existing backend Commander program. The backend package alias db:public-reader remains available to privileged maintenance wrappers. The backend entrypoint supplies its validated effective database URL, whether configured as a URL or separate connection settings; an explicit protected operator URL takes precedence. Both paths retain the exact reader target/context checks. Never supply credentials through command arguments, shell tracing, published plans or public-service configuration:
| Variable | Purpose |
|---|---|
WAVEMAP_PUBLIC_READER_ADMIN_URL | Optional explicit administrator URL; otherwise use the backend’s validated runtime connection. An explicitly empty value remains invalid. |
WAVEMAP_PUBLIC_READER_CONTEXT | Explicit target posture above. |
WAVEMAP_PUBLIC_READER_ROLE | Named dedicated reader. |
WAVEMAP_PUBLIC_READER_PASSWORD | Required for creation; optional explicit rotation for a retained role. At least sixteen characters, no NUL. Omit during ordinary reconciliation/checks. |
NODE_ENV | Exact test or deployed-development runtime posture. |
After migrations, with public traffic stopped and all reader sessions closed:
# Inspect the sanitized plan; no changes are applied.pnpm wavemap -- maintenance database public-reader
# Apply only after the target, role and proposed shared changes have been reviewed.pnpm wavemap -- maintenance database public-reader --execute --allow-shared-privilege-changes
# Remove the password from the operator environment before checking for drift.unset WAVEMAP_PUBLIC_READER_PASSWORDpnpm wavemap -- maintenance database public-reader --checkDefault invocation reports direct changes, inherited PUBLIC changes, potentially affected login roles and blockers. --check exits nonzero on blockers or drift. --execute changes nothing when blockers exist. The shared-change acknowledgement is needed only when the plan proposes database PUBLIC CREATE/TEMP revocation, public-schema CREATE revocation, unapproved public-routine execution revocation or restriction of the administrator’s future function defaults. The impact list is an inventory to review, not proof that every listed consumer is compatible. Unexpected defaults or inherited privileges outside that bounded contract are refused. New owners, migrations or manually changed grants require another review/check.
Password changes are transactionally coupled to grants and excluded from application reports and error diagnostics. Use an operator environment with appropriate PostgreSQL audit/log protection as well. On a connection/driver failure, commit status may be uncertain: keep traffic closed, inspect the plan and credentials, then retry deliberately. Do not infer rollback solely from a disconnected client.
Migration And Reset Integration
Section titled “Migration And Reset Integration”The existing privileged deploy:dev:db-migrate and deploy:dev:db-reset wrappers accept --public-reader-role <role> and optional --allow-shared-privilege-changes. Their default remains a local plan. Live --execute requires authority for that exact operation and target, using the existing operator SSM path; the ordinary application-delivery identity has no such authority. No new IAM permission is needed by this local implementation, and no permission is assumed granted to delivery automation.
With reader maintenance selected, the generated command checks retained role identity and closed sessions before schema mutation. This preflight intentionally does not require columns from pending migrations and does not prove final ACL readiness. It refuses a missing reader; first provisioning is a separate operation. The existing backend environment resolver supplies the same effective database connection used by migration/reset. The child environment removes any administrator-URL override and reader password so maintenance cannot target a different database or rotate credentials accidentally. No credential enters the generated SSM payload or process arguments.
Migration then reconciles the reader after pending migrations. Reset reconciles only after reset, migrations, base seed and Development seed all finish. The role and its password survive schema reset; object grants do not. These whole maintenance recipes are not one transaction: if a later stage fails, leave public admission closed and repair/reconcile before reopening. The owned maintenance recipe closes Caddy admission and stops the public container with a ten-second grace before preflight or SQL work. Reconciliation is followed by the full reader drift check. The session checks remain a second guard; they do not replace admission closure and drain.
Finish with a clean reader check, an actual rich public resource read, and main-app read/write smoke. Readiness alone only proves SELECT 1. The committed disposable tests cover retained identity/password, restored grants, forbidden access with read-only defaults disabled, and main authoring/fuzzy-search coexistence. Those tests do not establish live multi-consumer compatibility or launch capacity.
Development Seed Media
Section titled “Development Seed Media”Development import seeding admits only absolute HTTP(S) media URLs through the existing public URL contract. Legacy inline data: images and relative paths have no delivery asset and are omitted from ready media rows; the original import JSON remains intact. Such entities expose absent profile media until a real delivery URL is supplied. This seed-only repair does not relax response validation, rewrite existing deployed rows or claim that old third-party links remain reachable. Representative localhost image URLs target the separate local frontend and remain local demonstration data.
Isolated Local Container Proof
Section titled “Isolated Local Container Proof”apps/wavemap-public-api/Dockerfile.deploy prunes to the public app and its workspace dependencies, builds the real dist/bin/startServer.js entrypoint, installs production dependencies from the frozen pruned lockfile and copies only compiled workspace outputs and their dependencies into the runtime. It runs as the Node image’s non-root user. Backend sources, database maintenance tools, tests and application secrets are excluded.
The dedicated docker-compose.local.yml passes only the selected public settings. It publishes no host port and uses an isolated database network, a read-only root filesystem, dropped capabilities, no-new-privileges and an init process. Initial local controls are 0.5 CPU, 256 MiB memory, 64 PIDs, a 16 MiB temporary filesystem and ten-second stop grace. JSON logs rotate at 10 MiB with three files. These controls are local proof settings, not measured launch budgets or a guarantee against shared-host/database contention.
Build through the existing verifier, then explicitly load the selected image for the local runtime test:
pnpm verify:docker-builds -- --targets public-api-deploydocker buildx bake --file docker-bake.verify.hcl --builder wavemap-docker-verifier-buildkit-v1 --load public-api-deploy
# Use the name of the disposable PostgreSQL container bound to 127.0.0.1:55433.PUBLIC_API_TEST_DB_CONTAINER=your-disposable-postgres WAVEMAP_DB_TEST_TIMINGS=1 pnpm -F @wavemap/public-api test:containerThe test owns the existing guarded disposable baseline and must run sequentially with other database work. It validates the container’s published database port and identity, creates temporary networks, attaches only that database container, supplies a restricted reader URL through a private temporary environment file, and cleans up its containers, network attachment, fixtures, role and file. A disposable client container on that network exercises the actual listener with a rich Event response and a fifty-Event page with bounded Artist/Venue previews, HEAD, conditional 304, invalid-query response, the default seeded Event request and clean SIGTERM shutdown. Inspection verifies the effective user, command, resources, logs, networks, environment and absence of backend/source/socket mounts.
A responding mock metadata service lives on a separate network. The public container cannot reach it or the link-local metadata address in this local topology. This proves the local network configuration only: EC2 metadata isolation, instance-profile protection, actual ingress/egress rules and shared-host contention still require explicit deployment design and live activation evidence. The local Compose file is not the deployed three-image release contract. Its media proof uses seeded HTTP fallback URLs; optional S3/CDN-only locator configuration belongs to the later delivery configuration.
CI’s existing image-build lane selects public-api-deploy for public source changes and its shared dependency closure. It needs no private npm credential. The actual container test is an explicit local gate; it does not add another database CI job or duplicate the sequential database owner.
Deployed Release Capability
Section titled “Deployed Release Capability”The version-1 deployment contract has an optional but complete publicAPI capability: repository identity, provisioned or enabled mode, dedicated environment path, image-command parameter name and metadata-only reader-secret reference. The operations type derives from the validated projection. Absent capability retains the legacy two-image contract; partial capability is invalid.
In provisioned mode, the repository, third image and release metadata are modeled, but the host neither reads the reader secret nor starts a public service. enabled adds the isolated service to the existing Compose owner. An enabled host refuses old pair-only callers before rewriting environment or Compose files. Delivery and rollback retain the entire image set and activation posture; omitting an image cannot disable the service. Moving between provisioned and enabled states requires a reviewed operator transition and a new compatible baseline, because ordinary rollback refuses a posture mismatch.
The host reads only /wavemap/dev/runtime/public-api/PUBLIC_API_DATABASE_URL into the public environment file. It accepts the managed reader at db:5432/wavemap_dev, rejects URL overrides and writes a private file without exposing the credential in workflow metadata. Percent-encode reserved password characters. Public media URLs and region come from the modeled delivery configuration; no storage-write or main-backend credentials enter this file. Reader provisioning and secret population remain separate operator actions.
The enabled service uses the local proof’s resource/log limits, internal public_reader network and no host port. Only PostgreSQL also joins that network; a separate internal public_proxy network joins Caddy to the API. Readiness uses explicit IPv4 loopback because the listener binds IPv4. This is infrastructure configuration, not permission to enable public traffic: live HTTPS/origin trust, real EC2 metadata exclusion, maintenance admission closure/drain, live grants and rich read/denial proof must still be established. The maintenance gate closes admission and drains the reader around database work. Preparing an enabled source profile does not change the live deployment; its saved plan, contract publication, runtime delivery and controlled activation proof require the concrete operator checkpoint below.
Deployment Workflows owns full-range admission, whole-operation ownership, complete immutable receipts and rollback. An application image rollback does not undo a reader-policy change, migration, secret rotation or infrastructure transition.
Runtime source follows that same ownership: publicAPIRuntime.ts and publicAPIOrigin/runtime.ts select typed configuration, while adjacent .yml fragments own Compose syntax and bin/dev-deploy/utils/public-api-runtime.sh / public-api-origin.sh own host preparation. The shell support files define named helpers without doing work when loaded. Pulumi embeds them into the existing SSM recipe; they are not new CLI entrypoints. Only ${WAVEMAP_TEMPLATE_*} placeholders are bound during infrastructure rendering; ordinary shell variables remain for the existing host-side Compose heredoc.
Focused Release Lifecycle Proof
Section titled “Focused Release Lifecycle Proof”The opt-in Pulumi test below consumes an already loaded public image and postgres:16; it never pulls. It creates its own PostgreSQL volume and isolated networks, renders the actual public service/environment fragments, verifies readiness and effective restrictions, replaces the container and restores the previous image reference while retaining a database sentinel. Temporary aliases point to the same built image: this proves lifecycle/configuration behavior, not different application source revisions. It does not run migrations, use a shared database or replace the richer P4B ACL/resource proof.
WAVEMAP_TEST_PUBLIC_API_IMAGE=wavemap-public-api-deploy:latest pnpm -C infra/pulumi exec node --import tsx --test src/providers/aws/__tests__/publicAPIContainer.test.tsGenerated-host tests separately exercise failures before and after replacement, matching version restoration, old pair refusal and retained ownership. The real release-record wrapper runs against controlled AWS doubles to prove publication failure/retry and previous-receipt restoration. These are local tests; they do not establish live IAM, EC2 networking, CloudFront behavior or deployed capacity.
Development HTTPS Origin
Section titled “Development HTTPS Origin”The modeled request path is api.dev.wavemap.app → dedicated CloudFront distribution → Caddy on the existing EC2 host → private API container. The browser application keeps its own distribution and delivery behavior. The API adds no load balancer, second host or NAT gateway; Caddy shares the sleepable host and its resource budget. This avoids another always-on compute or load-balancer resource, but existing host/storage, CloudFront, DNS and traffic costs remain. Host headroom still needs measurement.
The existing Pulumi public-surface assembly owns the separate API viewer certificate in us-east-1, A/AAAA aliases and CloudFront policies. Both distributions reuse the dynamic app-origin.dev.wavemap.app record, so there is one owner for host-IP changes. publicAPIIngress provider outputs contain resource identities, paths and URLs, never secret values. These provider details do not broaden the neutral application release contract or create a new deployment CLI.
Provisioning And Request Policies
Section titled “Provisioning And Request Policies”Protected Pulumi config publicAPIOriginSecret opts into the origin resources. It must contain 43–128 base64url characters; supply a randomly generated value through the established protected configuration procedure. Optional publicAPIPreviousOriginSecret provides the rotation overlap slot. Missing current config retains the earlier runtime. Origin provisioning does not enable the reader: provisioned can start Caddy and prepare certificates while API requests receive a no-store JSON 503, with an empty HEAD body. enabled requires the origin and remains gated on maintenance, wake and live proof. Use the infrastructure change procedure and existing runtime deployment commands, with approval for the exact live action.
| Request | Edge Policy |
|---|---|
/v1/* reads | Every query parameter participates in cache identity; minimum/default TTL are zero and maximum TTL is 60 seconds. Origin cache headers determine actual freshness. Gzip/Brotli have normalized cache variants. |
/v1/health, /v1/ready, unmatched paths | No caching; query inputs still reach API validation. |
/.well-known/acme-challenge/* | No caching, no origin secret, no API provisioned gate, wake callback or activity callback. HTTP to the dedicated challenge listener is allowed. |
| API errors | No successful HTML rewrite; supported CloudFront error statuses have zero error-cache minimums. Origin no-store remains meaningful. Live cache/error behavior still needs edge proof. |
API requests forward the required Host, conditional and CORS headers plus CloudFront’s viewer address, without cookies or viewer authorization. Anonymous CORS permits GET/HEAD/OPTIONS and If-None-Match, exposes validators/retry/cache/request-ID headers, and allows no credentials. CloudFront admits only GET/HEAD/OPTIONS on the API origin-group behaviors. Unsupported methods are rejected by CloudFront with 403 Invalid Method; consumers must not expect the API’s JSON envelope for that service-generated error. In enabled mode, origin 502/503/504 responses select a private Function URL fallback. It requests wake once for GET/HEAD resource reads and immediately returns no-store JSON 503 with Retry-After: 60; HEAD is bodyless. OPTIONS, health/readiness, unmatched routes and certificate challenges do not request wake. The primary origin receives the public hostname for TLS; the signed fallback retains its own Function URL host. Local tests prove the modeled policy, while actual CloudFront failover and caching remain live activation checks.
Wake, Inactivity And Maintenance
Section titled “Wake, Inactivity And Maintenance”The wake, shutdown, inactivity, lifecycle and API fallback functions use the supported Node.js 24 Lambda runtime, matching the monorepo’s Node major. Their native modules need no separate transpilation or new dependency. Review the AWS runtime lifecycle before later activation or runtime changes.
The private lifecycle coordinator is the only automation that starts or stops the host. Reserved concurrency of one serializes its capability-specific aliases. Public callers can request wake only; the inactivity path can request stop only; delivery uses the private maintenance alias. One protected SSM String parameter outside the runtime-secret prefix persists automatic/maintenance mode, the maintenance owner and any pending start/stop intent. The host role explicitly denies reads of it, including through the otherwise broad AWS-managed SSM agent grant. This creates no database tables or end-user sessions. The host’s existing owner file and flock still protect whole operations and individual shell mutations.
The coordinator records intent before calling EC2. An uncertain response retains that intent until EC2 reports the corresponding transition; elapsed time never grants permission to reverse or forget it. Private status and same-transition retry support explicit recovery. Running-instance events and a fifteen-minute reconciliation schedule re-read the current EC2 address before updating the single shared origin A record. Delayed or duplicate DNS work cannot start the host. A missing address or failed update remains an error with bounded delivery retries.
The browser wake route returns immediately with 202 only when warming is accepted; a public API resource request remains 503 until the origin actually serves it. Consumers must implement bounded retry and honor Retry-After; ordinary fetch does not wait and retry automatically. Sixty seconds is a suggested retry interval, not a boot or certificate-readiness guarantee.
Each running application writes two container-local timestamp files through the shared server-only recorder: a heartbeat every twenty seconds and the last qualifying request, coalesced at five seconds. They contain no URL, viewer address, user identity or credentials. Registered public resource GET/HEAD reads, backend API requests excluding health/readiness/preflight, and frontend page requests excluding probes/assets/API paths count as activity. The frontend’s Node middleware runs the existing locale and auth chain as well as the recorder. The fixed SSM collector reads timestamps only; the apps receive no AWS credentials.
The inactivity monitor requires fresh evidence from frontend/backend and, when enabled, the public API. Every recorder must be fresh within ninety seconds. It allows five minutes after startup, then requests shutdown only when every service has been idle for 120 minutes plus a 65-second cache/write allowance. Invisible edge cache hits can extend the most recent origin read by at most the configured sixty-second cache lifetime. Missing, stale, future, malformed or failed evidence retains the host. This is a sampled idle decision: a request arriving between the final sample and stop can still encounter shutdown and retry through wake.
Maintenance acquisition closes public wake in the coordinator, privately wakes the host when needed, takes the existing operation owner, removes Caddy’s persistent open marker and drains the public reader. Caddy returns no-store 503 while the marker is absent, including after restart. Ordinary maintenance permits responses already fresh at the edge to finish their existing lifetime of at most sixty seconds. Urgent data removal requires separately authorized invalidation of affected cached paths; neither an application rollback nor origin closure purges edge cache.
Release requires frontend readiness, backend database readiness, a real rich Event read through the deployed public handler and zero-row private-access denial checks using its actual reader credential. The packaged verifier must exist in both the candidate and any rollback baseline. Cloud reopening and host-marker reopening are separate steps; interrupted release retains host admission closure and requires same-owner recovery. The exact procedure and failure caveat live in Whole-Operation Host Ownership. Never delete ownership or lifecycle state merely to unblock a failed run.
Certificates, Isolation And Forwarding
Section titled “Certificates, Isolation And Forwarding”Caddy uses the standard digest-pinned image selected in infra/pulumi/src/providers/aws/publicAPIOrigin/constants.ts; there is no DNS plugin or custom renewal worker. CloudFront uses an ACM viewer certificate. Caddy independently obtains a certificate for the public API hostname through Let’s Encrypt HTTP-01. The challenge passes through CloudFront over HTTP to host port 8080, allowing first issuance or expired-certificate recovery without working origin TLS. All other requests on that listener return 404. Ordinary API traffic uses host port 443 with the forwarded public hostname and distribution secret. See Caddy automatic HTTPS and CloudFront origin certificate matching.
The two listeners each have a CloudFront-only security group. The existing browser port-80 group remains separate because each CloudFront managed-prefix-list rule has weight 55 against a default 60-rule quota. Caddy also requires the distribution secret on ordinary HTTPS requests; the prefix list alone identifies CloudFront infrastructure, not this distribution. See AWS managed prefix-list weights.
The proxy runs as UID 1000 with a read-only root, dropped capabilities except NET_BIND_SERVICE, no-new-privileges, 0.25 CPU, 128 MiB and bounded logs/PIDs. Its egress bridge permits ACME access. Its private API bridge is separate from the reader/database bridge; the API has no internet-facing bridge or published port. Caddy strips caller-supplied forwarding headers and replaces the API’s internal viewer-address header from CloudFront-Viewer-Address. The API trusts that value only from the configured fixed proxy address, validates its IP and source port, then applies existing limiter normalization. Other peers use their transport address. This is request provenance for an anonymous API, not account authorization. Cache hits do not reach the process limiter.
The private proxy network reserves separate address space for dynamic reader allocation. PUBLIC_API_PROXY_DYNAMIC_RANGE and the fixed proxy address are owned by publicAPIOrigin/constants.ts and rendered through the native networks.yml fragment. Docker can restore the reader before Caddy during host boot; allowing its dynamic pool to include Caddy’s fixed address can prevent the proxy from starting. Startup order is therefore insufficient protection. Reader-first startup and daemon-restoration tests exercise this boundary.
Before starting public containers, the existing host recipe installs a narrowly scoped metadata guard for the three dedicated public bridges. It adds DOCKER-USER denies for 169.254.169.254, leaving the main backend’s instance-profile access intact. IPv6 is disabled on those networks. A Docker ExecStartPre drop-in reinstalls the rules before container restoration, and deployment checks the running daemon’s forwarding hook without restarting Docker. Missing iptables or an unsupported forwarding posture fails closed. The implementation follows Docker’s user-chain boundary; actual Amazon Linux boot/systemd and EC2 metadata denial remain live verification gates.
Recovery And Secret Rotation
Section titled “Recovery And Secret Rotation”The host writes /opt/wavemap/public-api.Caddyfile and the proxy-only /opt/wavemap/env/public-api-origin.env atomically. The latter is mode 600 and contains only the current/overlap origin credentials. A configuration fingerprint in Compose triggers proxy replacement when the mounted Caddyfile changes. Application rollback does not restore older infrastructure configuration or credentials.
Certificate/account state lives in /opt/wavemap/public-api-certificates, mode 700 and owned by UID 1000. The API never mounts it. Container replacement retains this directory; EC2 replacement does not guarantee retention because the current root volume is deleted on termination. Preserve the directory during ordinary repair. Losing it requires fresh issuance and can encounter public CA rate limits.
Changing Compose network allocation does not update an existing Docker network in place. A separately approved transition must retain maintenance ownership and closed admission, validate the exact candidate with tools supported by the host, preserve the original Compose file, and replace only the affected public containers and private proxy network. Keep the existing images, database and certificate mounts. Verify the actual TLS listener and origin-credential denials alongside rich reads, retained data and metadata isolation before the guarded release. The application maintenance verifier alone does not prove that Caddy is reachable. A full stop/wake drill requires its own bounded approval and must demonstrate public recovery without manual container ordering.
Caddy renews only while running. After a long sleep, first restore the dynamic origin DNS record and HTTP challenge route, then allow Caddy to recover TLS before treating the API as available. Do not delete certificate storage or repeatedly recreate the service to troubleshoot a CA/DNS outage. Inspect bounded proxy logs and challenge reachability first. There is no periodic wake solely for certificate renewal.
For Development, the owner accepted real private-CA renewal, CA-outage retention and expiry-recovery tests together with live public issuance, retained-certificate restart and changed-IP wake recovery on September 19, 2026. This acceptance preserves the deployed certificate rather than deleting it or changing the host clock to force an expiry drill. It does not establish natural public-CA renewal, public CA/DNS outage recovery or long-sleep public-certificate expiry recovery; those remain production-readiness gates. Origin-secret overlap and retirement still require their separate live proof. Revisit this acceptance if the proxy image, certificate storage, challenge routing or wake/DNS behavior changes.
Rotate credentials through overlapping acceptance, using separately reviewed infrastructure and host deployments:
- Keep the old current secret. Put the new candidate in
publicAPIPreviousOriginSecret, apply the reviewed configuration and deploy the host so Caddy accepts both. Despite the key’s name, this slot can temporarily hold the next value. - Set
publicAPIOriginSecretto the new value and the overlap slot to the old value. Apply the reviewed infrastructure change and wait for CloudFront deployment completion. The running proxy already accepts both values. - Remove the overlap config, apply, and deploy the host to retire the old credential. Confirm ordinary HTTPS succeeds with the new value and rejects the old value. Never send credentials over the challenge route or include them in reports.
If the metadata guard prevents Docker from starting, repair the supported firewall/guard installation through the operator path while public admission remains closed; bypassing the prerequisite can expose instance credentials. Removing origin config does not automatically remove the installed host drop-in, guard or persisted certificates. Retiring those files/rules is an explicit host cleanup after the public containers are stopped.
Permission Prerequisites
Section titled “Permission Prerequisites”These are source-backed requirements, not verified live grants. Follow Cloud Automation Permission Review before applying:
| Actor | Required Scope And Boundary |
|---|---|
| Infrastructure operator | Create/update/read/delete the named API CloudFront distribution, policies and provisioned function; request/describe/delete/tag its ACM certificate in us-east-1; change/read the API aliases and validation records in the existing Route53 zone; create/tag/manage the two security groups and attach them to the existing instance; inspect the managed prefix list; put/get/delete the exact origin SecureStrings. Include provider wait/readback/tag operations. The operator policy is externally managed and must be reviewed before live apply. |
| Existing EC2 instance role | Read /wavemap/dev/runtime/public-api-origin/PUBLIC_API_ORIGIN_SECRET and the optional PUBLIC_API_PREVIOUS_ORIGIN_SECRET with decryption. Its existing runtime-prefix SSM grant covers these paths. Parameters use the default AWS-managed SSM key; selecting a customer key requires a separate kms:Decrypt and key-policy review. |
| Existing delivery/operator path | Invoke the established host deployment document within its current authority. Application delivery also invokes the exact private maintenance alias, with no secret read, Route53, certificate, direct EC2 or Pulumi authority. Privileged database maintenance remains separate. |
| Caddy/API containers | No AWS identity or DNS credentials. Only the host retrieves the proxy secret; only the reader receives its database credential. Caddy calls the public ACME service. |
Permission Readback And Staged Activation
Section titled “Permission Readback And Staged Activation”Before applying, compare the intended source policies with the effective live actors. Read-only inspection does not grant missing authority. Store raw IAM/trust/resource-policy documents, simulator results and provider identifiers in private operator evidence; publish only the decision and remaining gaps. An expired SSO session or denied read is a blocker, not proof of an absent policy.
- Confirm the configured account/region and infrastructure-operator identity with STS. Inspect its permission set/role policies, boundary and organization controls for the named CloudFront distribution/policies/functions/OAC, ACM certificate, Route53 records, security groups, SSM parameters/documents, Lambda aliases/event rules/log groups and exact IAM roles. Include provider read/tag/wait/cleanup operations and PassRole only to the intended services. Monitoring also needs CloudFront Create/Get/DeleteMonitoringSubscription on the API distribution, CloudWatch alarm management on the named alarms, log metric-filter management on the existing groups and SNS subscription management on the operational topic; an email subscription requires recipient confirmation. The operator remains externally managed.
- Inspect application-admission and application-delivery role trust for the exact repository, protected environment subject and STS audience. Read all attached/inline policies and any boundary. Admission remains metadata-only. Delivery requires the complete image/release set, exact deployment document/instance and private maintenance alias; it must not gain AWS-RunShellScript, reader-secret access or direct lifecycle mutation. Inspect GitHub environment protections and the retained reader-role variable separately.
- Inspect the privileged operator’s exact ECR/release-publication, host document, database-maintenance and maintenance-alias authority. Check the EC2 instance-profile role, modeled image pulls, dedicated runtime parameter reads, lifecycle-state deny and existing version/media/SSM grants. No permission simulator result proves database ACLs or container metadata isolation.
- After the approved apply, inspect each Lambda execution role and resource policy, OAC and qualified alias target. Confirm the coordinator’s single concurrency, protected state, exact DNS conditions, fixed activity document and restricted invoke sources. Verify alert destinations and confirmed subscriptions; an alarm with no working recipient is not notification delivery.
Use a clean exact source revision for the refresh-enabled full-stack saved plan. Review plan/source digests, IAM/trust changes, persistent lifecycle state, document/configuration changes and all replacements/deletions before requesting the exact apply. Reader provisioning, secret population, contract publication and runtime/database operations each need their own concrete target and approval. Keep public mode provisioned while establishing a compatible three-image candidate and rollback baseline, retained reader, origin credentials and host maintenance capability.
The live proof record should identify the source SHA, image digests, contract version, approved plan and each observed result. Record the following evidence and any explicit environment-specific acceptance before treating activation proof as complete:
| Proof | Required Observation |
|---|---|
| Delivery and recovery | Exact-source three-image deploy and application rollback preserve the database sentinel; a failed/interrupted operation retains admission closure and resumes with the same owner. Both images include the maintenance verifier. |
| Reader and coexistence | Real rich resource reads succeed; private fields/tables and mutation are denied under the deployed reader; main-app read/write and fuzzy-search behavior still work after regrant. |
| Edge and browser | HTTPS, origin-bypass denial, query identity, malformed-after-warm requests, CORS-visible retry/validator headers, HEAD/304/compression and zero sticky error/stale-success behavior pass through actual CloudFront. |
| Lifecycle | A sleeping origin gives prompt truthful retry metadata, then a real read; stopped/pending/stopping/unhealthy and maintenance cases remain distinct. Probe-only traffic permits idle stop, API-only traffic retains the host, and absent evidence prevents stop. Measure fallback delay and cache allowance. |
| DNS and certificates | Changed-IP wake repairs the shared record; duplicate/missed completion recovers. Prove public issuance and origin-secret overlap/retirement without deleting retained certificate state. Public renewal, CA/DNS outage and long-sleep expiry proof remain production gates under the Development acceptance above. |
| Isolation and notifications | Public containers cannot reach real EC2 metadata before/after Docker and host restart while the main backend retains its intended access. Confirm alarm/budget recipients and a separately approved test notification. |
A failed case leaves admission closed or mode provisioned. Repair under the same owner; use private status, recorded-transition retry and DNS reconciliation where applicable. Do not erase uncertain lifecycle intent, bypass metadata protection, reset data as rollback, or enable traffic solely because readiness passes. P5 owns representative contention and cost measurements; P6 owns consumer documentation and launch.
Focused Origin Proof
Section titled “Focused Origin Proof”The opt-in tests require the exact image references already loaded locally and never pull them. Read the test constants before loading images. They own disposable containers/networks/storage and do not mutate AWS, issue public certificates or change the developer’s firewall. The metadata lane uses an isolated privileged nested Docker daemon, without a host socket, to test real forwarding rules and daemon restoration with a responding positive control.
# Roughly six minutes: real short-lived certificate renewal and expiry, using a private CA.WAVEMAP_TEST_PUBLIC_API_ORIGIN=1 pnpm -C infra/pulumi exec node --import tsx --test src/providers/aws/__tests__/publicAPIOrigin.test.ts src/providers/aws/__tests__/publicAPIOriginContainer.test.ts
# Roughly one minute: actual packets before and after nested Docker restart.WAVEMAP_TEST_PUBLIC_API_FIREWALL=1 pnpm -C infra/pulumi exec node --import tsx --test src/providers/aws/__tests__/publicAPIOriginMetadata.test.ts
# Roughly one minute: generated Caddy/API services and a fresh PostgreSQL database.WAVEMAP_TEST_PUBLIC_API_ORIGIN_IMAGE=wavemap-public-api-deploy:latest pnpm -C infra/pulumi exec node --import tsx --test src/providers/aws/__tests__/publicAPIOriginIntegrated.test.tsThe integrated lane verifies TLS/hostname, readiness, validation, CORS, limiter identity, network/credential separation, API replacement and a configuration-driven proxy replacement with the CA offline. It checks certificate and database-sentinel retention. Its database proof is readiness/lifecycle, not migrated resource queries or ACL reconciliation; retain the separate rich-reader and denial tests above. Pulumi resource mocks separately check the provisioned/enabled edge model. These tests do not establish live CloudFront caching/CORS, IAM, EC2 isolation, public issuance or measured capacity.
Operating View
Section titled “Operating View”Start with one selected time window and these existing AWS/host surfaces. A stopped Development host is expected; an intentionally closed maintenance gate is also expected. Neither should be treated as an always-on availability incident. Keep the API distribution distinct from the browser distribution when comparing requests or failures.
| Signal | Read It From | Interpretation And Response |
|---|---|---|
| Traffic and transfer | API distribution Monitoring: Requests Sum and BytesDownloaded Sum in AWS/CloudFront, region us-east-1 and dimension Region=Global. | Includes probes, cache hits and errors; never use this total as the filtered inactivity signal. Compare a burst with transfer, budget and fallback attempts. |
| Cache effectiveness | CacheHitRate Average on that distribution, when additional metrics are enabled. | This is the aggregate cacheable-request hit rate. Missing data means unavailable, not zero. A repeated warm test proves that URL only and must not be presented as fleet-wide effectiveness. |
| Latency | OriginLatency p95 when additional metrics are enabled; public completion durationMs in bounded container logs. | Origin first-byte latency and application completion duration measure different boundaries. Compare equivalent windows; small samples and missing periods limit conclusions. |
| Errors and throttling | Default CloudFront 4xxErrorRate/5xxErrorRate; public completion status and errorType. | CloudFront’s 4xx rate is not a separate 429 count. Origin logs distinguish rate_limited and unavailable but omit edge hits/failures. Check intentional maintenance and wake before treating a 503 as application failure. |
| Wake and DNS | Fallback log source wavemap-public-api-fallback with accepted/refused/failed outcome; lifecycle source wavemap-runtime-lifecycle with capability/status; native Lambda Errors/Throttles. | Accepted means start accepted or already pending/running, never API readiness. A lifecycle starting record identifies a new accepted start; dns/failed identifies DNS completion failure. Inspect private status before an explicit same-transition retry or DNS reconciliation. |
| Inactivity | Seven-day monitor logs: action/reason, plus its fixed readback result. | Repeated stale/missing/unavailable evidence retains the host and can waste money. Repair recorders/SSM permissions; do not reinterpret missing evidence as idle. Healthy probes should leave activity unchanged. |
| Process and pool | Public completion runtime fields: accepting, activeReads, pendingStatements, poolMax, maxActiveReads and residentMemoryBytes; host docker stats —no-stream for all Compose services. | Gauges are snapshots at completion, not historical peaks. Pending statements include queued driver work; they are not a PostgreSQL connection count. Compare API and main-app/container pressure together. |
| Database pressure | Privileged read-only aggregate pg_stat_activity snapshot below. | Compare client/reader connections, active sessions and lock waits with max_connections. Do not print query text, connection strings, users or client addresses. This is an operator diagnostic, not a new public endpoint or reader grant. |
| Host capacity and spending | Existing disk/inode alarms and the account’s Development monthly budget. | Capacity alarms are 80% for two thirty-minute periods. The USD 30 monthly budget notifies at 50%/80% actual and 100% forecast. Neither is a spending cap or per-API cost attribution. |
CloudFront provides traffic and aggregate error metrics by default. The modeled API distribution enables its additional-metrics subscription for aggregate CacheHitRate and OriginLatency; AWS sells this as a bundle of up to eight charged metrics. At the first US East metric tier, allow approximately USD 2.40/month for that bundle, USD 0.30 for one shared control-failure metric and USD 0.30 for three standard alarms: approximately USD 3/month incremental before retrieval/notification charges, taxes and any free-tier effects. This is an estimate, not a spending cap. The owner approved this monitoring budget; the exact infrastructure apply remains a separate action. There is no paid dashboard or new log-shipping pipeline. Check CloudFront metrics and CloudWatch pricing when changing that choice.
Under the existing privileged operator access, this aggregate-only query gives a bounded database snapshot. It performs no schema/data mutation and needs no new API credential:
SELECT count(*) FILTER (WHERE backend_type = 'client backend') AS client_connections, count(*) FILTER (WHERE backend_type = 'client backend' AND state = 'active') AS active_sessions, count(*) FILTER (WHERE wait_event_type = 'Lock') AS lock_waiters, count(*) FILTER (WHERE application_name = 'wavemap-public-api') AS public_reader_connections, current_setting('max_connections')::integer AS connection_limitFROM pg_stat_activityWHERE datname = current_database();Retention, Alerts And Recovery
Section titled “Retention, Alerts And Recovery”Public completion logs contain generated request IDs, registered route labels, bounded HTTP/error fields and allowlisted numeric runtime gauges. They exclude raw URLs/queries, public entity IDs, client addresses, headers, SQL and credentials. Do not turn request IDs into metric dimensions. Container logs rotate at 10 MiB times three files per service; this is a size cap, not a guaranteed number of days. Read bounded windows with Compose logs —since/—tail and keep raw output private. Lambda log groups retain seven days; SDK/platform diagnostics can still contain infrastructure identifiers. There is no new centralized application-log store.
The disk/inode and lifecycle alarms use the existing regional operational SNS topic. Its modeled email subscription uses the separate protected operationalAlertEmail Pulumi setting; applying it sends a confirmation request, and the recipient must confirm before alarms can be delivered. The current Development recipient is the Tendril management account root email, chosen as an interim destination. Replace this setting when a dedicated operations mailbox is available; confirm the replacement SNS subscription before retiring the old one. Budget notifications retain their separate budgetAlertEmail destination. Confirm the live subscription state and separately authorize a test notification before relying on delivery. A topic with no confirmed recipient is incomplete.
Three lifecycle alarms evaluate two of three fifteen-minute periods and send both alarm and recovery transitions. The shared control-failure count detects unresolved lifecycle recovery, failed API wake calls and missing/invalid/stale inactivity evidence. Native coordinator Errors detects exceptions and timeouts, including DNS failure; native Throttles detects at least five rejected concurrent invocations per breaching period. This higher throttle threshold tolerates occasional contention without weakening the coordinator’s required single concurrency. All three alarms use the workload region and existing topic; CloudFront metric views use us-east-1.
The three log filters publish one dimensionless metric per deployment prefix, so request IDs, owners and outcome labels cannot multiply the metric bill. Normal maintenance, accepted wake, startup grace, recent activity and a stopped host do not increment it. Native Errors covers coordinator failures that throw, while its log filter covers non-throwing recovery decisions. These counts can include retries; inspect the bounded source logs to find the cause. Missing data is treated as non-breaching because this is a sleepable dev runtime. Consequently these alarms do not establish scheduler/log-delivery health, continuous API availability or historical process/database peaks; retain the operating checks above.
For capacity pressure, retain maintenance ownership while investigating disk/inodes, container usage and PostgreSQL connections. Use the existing image cleanup policy; never delete database or certificate storage as a quick fix. For failed wake/DNS or repeated inactivity-readback failures, inspect private lifecycle status and the exact role/document permissions before retrying. For a budget or abusive-traffic incident, close public admission through the owned maintenance path and review edge invalidation/configuration separately; the process limiter does not cap cache-hit or fallback costs. Existing fresh cached responses may remain visible for sixty seconds. The maintenance procedure owns recovery and verified reopening.
Operator Response Checklist
Section titled “Operator Response Checklist”Use the existing maintenance owner and deployment tools for live changes; prepare the exact target and mutation plan before requesting approval. Start with private lifecycle status so expected sleep or maintenance is not confused with overload. Do not open admission merely because a health check passes.
| Trigger | Inspect First | Bounded Response | Verify Before Reopening |
|---|---|---|---|
| Cache or query-contract change | Effective CloudFront cache policy, query-key forwarding, origin cache headers and the generated v1 artifact. | Preserve all representation-varying controls in the cache key; errors and maintenance remain no-store. Treat TTL changes and invalidation as reviewed deployment actions. | Distinct valid query variants do not share bodies; warm GET and HEAD retain the right validators; 304 is bodyless; 400, 429 and 503 retain their status and retry/cache metadata. |
| Cancellation, removal or private-data exposure | Whether the change is a public cancellation, lifecycle removal, changed relationship or urgent privacy incident. | Apply the domain-owned writer. Ordinary cached success can remain for up to sixty seconds. For urgent removal, close admission under ownership and separately approve invalidation of every affected cached representation; rollback does not purge caches. | Fresh origin reads and affected edge variants reflect the intended visibility; cancellation remains readable where the contract requires it; removed details are absent and complete collection reconciliation excludes removed membership. |
Sustained 429 or 503 | Rate-limited versus unavailable log labels, filtered activity, reader admission, connections, lock waits, container memory and the main application’s health. | Stop test traffic first. Under the approved incident scope, close public admission while diagnosing. Keep the configured pool/timeouts; do not raise limits until shared-host headroom is measured. | Pending work drains, connections remain bounded, a real public read succeeds and the main app recovers; release the same owner through its verified path. |
| Cost spike or abusive traffic | Distribution request/transfer totals, cache effectiveness, fallback invocations, running hours and account-wide budget usage. | Origin admission protects origin work, but cache-hit delivery and fallback requests can still cost money. Prepare separately approved edge restriction/disablement if needed; a process limiter is not a spending cap. | Confirm the selected control affects the costly traffic path, preserves intentional main-app access and has a recorded reversal. Check usage after billing/metric delay. |
| Bad release | Last successful immutable receipt, current images and schema compatibility. | Use the guarded runtime rollback with retained images and the existing owner; never reset the database as rollback. | Reader grants/private-data denial, metadata isolation, real API reads, main-app behavior and open/automatic/unowned lifecycle state. |
Measuring Capacity Without Inventing Guarantees
Section titled “Measuring Capacity Without Inventing Guarantees”Record the source/image identity, effective limits, dataset shape, client location, request set and sampling window before a workload. Label a new valid query key as an edge miss only after observing the edge response; do not claim the database cache is cold without proving that independently. Measure repeated keys, varied valid keys, simultaneous misses, maximum includes, cursor traversal, conditional reads and incremental refresh separately. Record successes, overload responses and failures as separate populations; a fast 503 is not a successful-query latency improvement.
Use the operating view above for latency, response bytes, resident memory, reader connections and main-app impact. Completion gauges and periodic snapshots do not prove an unobserved peak. HEAD and origin conditional reads still perform GET’s database work. Shared infrastructure means a passing API-only run cannot establish safety under simultaneous application load. Keep raw request evidence private and publish only bounded summaries and limitations.
Stop a workload on unexpected ownership, main-app failure, sustained database/host pressure or exhausted request/byte/time caps. Keep initial capacity values provisional until the repository owner reviews representative results. The existing defaults and container limits are enforcement settings, not measured sustainable throughput.
Running The Bounded P5 Workload
Section titled “Running The Bounded P5 Workload”Use the capacity command to print the exact request plan without contacting either service:
pnpm wavemap -- smoke dev public-api-capacityThe plan has 190 GETs: thirty repeated reads, thirty valid query variants, eight simultaneous pairs, fifty rich/traversal slots, twenty conditional slots (ten initial/validator pairs), twenty incremental/traversal slots and twenty-four main-app reads interleaved across the run. The incremental sample starts at the documented fixed historical lower bound; it measures the supported query path, not a real consumer’s freshness or deletion detection. Server continuations replace slots and remain on the same origin and route. Exhausted collections restart at the root for warm samples. A small dataset can therefore produce few continuation observations; report that limitation. A pair is only a simultaneous-miss sample when its recorded cache headers establish that fact.
Execution requires curl 8.4 or later, a new private output directory and a private file containing an existing main-app bearer token. That token is solely for comparing authenticated Artist reads on the main application. Public API requests remain anonymous and receive no authorization, cookies or viewer-address header. Obtain the token through the existing main-app login flow; do not put it in command arguments or source control. The token file must have no group/other permissions (chmod 600). After the inactivity trial is complete and the concrete live checkpoint is approved:
pnpm wavemap -- smoke dev public-api-capacity \ --output-directory "$PWD/.local/public-api-p5-run" \ --main-bearer-file "$PWD/.local/main-app-token" \ --executeOnly the fixed Development origins or numeric HTTP loopback origins are accepted; both targets must be in the same environment. There are no redirects or retries. Dispatch averages at most one request per second, with concurrency two only in the eight selected pairs. The runner stops at 240 requests, a 250 MB body budget or twenty minutes, with a twenty-second per-request timeout and a 16 MB response stop threshold. It reserves 1 MB per in-flight request for curl’s final receive buffer and accounts for actual temporary-body bytes when curl reports zero after a limit error. These are client stop thresholds, not a network firewall or a billing ceiling. An oversized response, transport failure, unexpected status, main-app failure, invalid continuation or three consecutive 429/503 responses stops the run and retains partial evidence.
Ctrl-C cancels active curl requests. Creating STOP inside the output directory prevents the next batch; it can take the current request timeout to take effect. The runner prints phase progress and writes a private plan, a journal after every batch, and a final summary. Request numbers connect observations to plan slots even when pair completion order differs. Summaries separate route, phase, HTTP status and observed cache header, with sample counts, bytes and latency percentiles. Bodies, cookies and bearer credentials are absent from saved reports. Native curl opens a fresh connection for each request and requests no compression, so timings include that client’s connection costs and body sizes are not a substitute for billed compressed transfer. Capture the actual CloudFront billing/transfer view separately when completing the cost worksheet.
This HTTP tool does not monitor AWS ownership, Docker or PostgreSQL. Run it with an operator observing the existing lifecycle and runtime controls. Before starting, record the deployed image/source receipt, effective runtime limits, dataset counts, CloudFront configuration and healthy automatic/unowned posture. Capture aggregate container CPU/memory, restart/OOM state and PostgreSQL connections/lock waits at start, middle and end, alongside the journal’s main-app samples. Use the existing operating view and access; keep query text, credentials and application rows out of evidence. Stop immediately with Ctrl-C on unexpected ownership, a restart/OOM or persistent lock pressure. Record unavailable observations explicitly and do not accept capacity on incomplete host/database evidence. The finite workload does not establish sustained contention, a maximum throughput ceiling or production availability.
Cost Worksheet
Section titled “Cost Worksheet”Estimate edge requests and delivered bytes separately from origin misses: requests / 10,000 × request price + delivered GB × transfer price. Cache hits reduce origin work but still serve bytes. Add CloudFront Function invocations, fallback/coordinator Lambda requests and GB-seconds, extra EC2 running hours attributable to API activity, public IPv4 hours, storage, DNS, metrics/alarms, log ingestion/storage and any approved invalidations. Shared host/storage costs should not be charged twice; record incremental and total-environment estimates separately.
For a planning example dated September 19, 2026, assume thirty days, North American HTTPS delivery, 0.00005 billed GB per normal response or 0.00025 GB per large response, and no free-tier/discount allocation. AWS lists USD 0.010 per 10,000 HTTPS requests and USD 0.085/GB in the first paid transfer tier. The following are gross edge-only estimates, not traffic measurements or total bills. Verify rates and shared free-tier allocation in CloudFront pay-as-you-go pricing before a release decision.
| Assumed Traffic | Monthly Requests | Monthly Transfer | Gross Edge Requests And Transfer |
|---|---|---|---|
| 10,000 requests/day at 0.00005 GB | 300,000 | 15 GB | About USD 1.58 |
| 1,000,000 requests/day at 0.00025 GB | 30,000,000 | 7,500 GB | About USD 667.50 |
The high-volume case is a sensitivity scenario, not an instruction to load the Development endpoint. Use actual compressed response sizes and geographic delivery mix when replacing these assumptions. The existing roughly USD 3/month monitoring choice and USD 30 Development account budget remain separate; budget notifications do not cap spending and confirmed alert delivery is still a Development deferral. Sustained origin traffic can also keep the shared host running all day. Use EC2 pricing and the actual instance/region to price the additional hours rather than assuming that the API’s existing host is free.
CloudFront also offers flat-rate plans. This worksheet assumes usage-based billing; it does not assume that this distribution is enrolled in a plan. A plan change requires checking its feature compatibility and separately reviewing the infrastructure and cost tradeoff.
Maintenance Ownership
Section titled “Maintenance Ownership”Update this page when changing the backend reader contract/reconciler, typed migration/reset planners, target guards, public runtime, publicAPIOrigin or publicAPIEdge configuration. See Public API Foundation for query and process ownership, Runbooks for existing database operations, and Cloud Automation Permission Review before introducing a new execution identity.