Skip to content

Infrastructure Change Policy

Infrastructure changes are graph transitions, not just code changes. A passing typecheck can still describe replacement, deletion, new trust boundaries, new costs, or data loss.

This page records the current shared-environment policy for Pulumi-managed infrastructure and cloud automation.

Pulumi owns the shared cloud resource graph:

  • Resource names, tags, IAM roles, policies, hosted zones, certificates, buckets, distributions, runtime host, control plane, deploy documents, and stable outputs.
  • Provider-specific implementation code under infra/pulumi/src/providers/aws.
  • Provider-neutral output contracts where those concepts are real, such as registry, ingress, runtime host, media storage, runtime configuration, deployment target, and control plane facts.

The delivery and operator workflows own app release orchestration:

  • SHA-tagged image identity.
  • Backend/frontend image build and push.
  • Runtime config validation.
  • SSM deploy invocation.
  • Reset/seed controls.
  • Smoke checks.
  • Runtime release receipts and app-runtime rollback.
  • Telemetry summaries.

Do not let application delivery or a deployed-dev operator recipe quietly become infrastructure mutation.

Use lifecycle classes when reviewing Pulumi previews, topology projections, runbooks, and deployment summaries.

ClassMeaningCurrent ExamplesCD Posture
ephemeralCan be recreated without preserving identity or data.Lambda code packages, generated cold-start object, short-retention logs.May be updated automatically after the workflow and preview policy are trusted.
replaceableMay be replaced, but dependent outputs or runtime facts must be refreshed.Disposable EC2 runtime host posture, dynamic origin targets.Requires visible preview summary and follow-up readiness checks.
persistentHolds data, identity, policy, budget, DNS, or history that should not be casually lost.Media bucket, hosted zones, budgets, runtime root volume, IAM trust boundaries.Requires explicit human approval for replacement/delete and may need retain/protect/runbook work.
externally durableDurability is provided by another managed service or external system, but topology changes still need care.Future managed database or separately mounted data volume.Treat changes as migrations with validation, not routine updates.

When in doubt, classify the resource conservatively and ask what would be lost if the graph transition applied exactly as previewed.

Application and docs delivery must not run pulumi up.

Guarded infrastructure changes should:

  • Use bin/infra/pulumi-stack.sh from a clean exact source revision.
  • Create one refresh-enabled, full-stack saved plan under a private ignored path.
  • Review creates, updates, replacements, deletes, protected resources, and IAM/trust-boundary changes.
  • Call out persistent or data-bearing resources explicitly.
  • Include the expected operator action for replacements, deletes, DNS delegation, SSM/runtime impact, and rollback.
  • Apply only the exact reviewed plan after explicit approval, supplying both its SHA-256 and source Git SHA.

Special handling is required for:

  • Stateful replacement or deletion.
  • Database or storage topology changes.
  • DNS delegation and certificate changes.
  • IAM broadening or OIDC trust changes.
  • Runtime host replacement.
  • Media bucket replacement or destroy.
  • Anything that can orphan, hide, or delete data.

Pulumi previews should be reviewed as graph-transition plans. They are not just “does the infra package compile?” evidence.

When a graph transition adds automation authority, also use Cloud Automation Permission Review to name the actor, cloud API actions, resource scope, trust boundary, evidence, and explicit non-goals before the automation is treated as ready to run.

The deployed dev app currently relies on CloudFront as the public ingress boundary, private S3 origins through Origin Access Control, and EC2 app ingress restricted to the AWS-managed CloudFront origin-facing prefix list.

AWS WAF/Web ACLs are intentionally deferred for deployed dev while traffic is low and cost minimization remains the priority. Revisit this posture when request volume, anomalous wake traffic, spam, or abuse evidence justifies the recurring cost and operational surface.

For public-facing deployments such as production, this protection should be treated as a stronger baseline candidate rather than a last-resort add-on. A production review should consider an ingress Web ACL on CloudFront with rate-based rules and managed rule groups sized to normal traffic before the surface is opened broadly.

Pulumi State And Deployment Contract Boundary

Section titled “Pulumi State And Deployment Contract Boundary”

Pulumi state is infrastructure-control-plane state. It records the full resource graph, provider identifiers, config encryption metadata, backend locking facts, and other implementation details needed by infra operators or approved infra automation. Application and docs delivery workflows must not treat direct Pulumi state access as a deploy input.

Deployment contracts are the sanitized deploy-facing artifacts derived from approved stack outputs. They should contain only the fields app/docs deployment needs: target, region, public URLs, registry references, runtime host paths, runtime configuration references, media/docs delivery facts, explicit deploy executor handles, and output-contract metadata. They should not include raw Pulumi exports, arbitrary provider physical IDs beyond explicit deploy handles, secret values, or state backend implementation details. The initial provider-neutral contract is wavemap.deployment-contract schema version 1.

The active backend is the Tendril Systems state foundation. Tendril owns the protected backend resource and its organization-level controls; Wavemap uses only the dedicated /wavemap namespace for the product stack. Wavemap still owns the wavemap-infra project, aws-dev stack, workload resources, passphrase, guarded launcher, and recovery procedure. Sharing the foundation does not give Wavemap authority over /foundation or /waveguide state.

Wavemap’s former workload-account S3 backend is retained as protected recovery evidence. It is not the active backend, not a deployment input, and not a pending migration target. Exact object names, account identifiers, checkpoint exports, and recovery material remain in private operator evidence.

The deployment contract store is independent from both backends. The first contract-store interface is TDeploymentContractStoreAdapter; it supports local-file no-cloud workflows and AWS S3 as the first private cloud-backed store. Future Azure Blob, GCS, or other private artifact stores should be added only as real adapters when needed.

The AWS dev stack now provisions the first private deployment contract artifact store as an S3-backed aws-s3 store. Its current-contract objects live under deployment-contracts/<pulumiStack>/current.json, are stored in a non-public bucket with bucket-owner-enforced ownership, default SSE-S3 encryption, and S3 versioning, and are exported from Pulumi as deploymentContractStore so follow-up automation can configure the storage adapter without knowing raw provider state.

This artifact store is not the Pulumi state backend. It is the deploy-facing storage surface for sanitized contracts after an approved infra pass has produced stack outputs. The Pulumi-managed docs deploy role has read-only access to the deployment contract prefix, and docs publication reads that contract instead of direct Pulumi outputs. The app deploy contract projection now includes runtimeDeployment.executor for the app lane, currently shaped as an aws-ssm-send-command handle with document name and target instance ID. The app deploy contract-read grant is a separate Pulumi-managed policy attachment to the GitHub application-delivery role, scoped to list/read the contract prefix only. Application and docs delivery use this sanitized projection rather than Pulumi credentials or backend access. Any future automated contract publisher still needs a separate least-privilege writer review.

The durable deployed-dev operator model separates state, sanitized deploy facts, runtime secrets, and workflow bootstrap values. Reaching for Pulumi state, GitHub secrets, or raw provider outputs during delivery is a signal that the deploy contract boundary needs to be refreshed.

SurfaceCurrent LocationWritersReadersMust Not Own
Pulumi state backendTendril state foundation, dedicated /wavemap namespace.Guarded Wavemap infra operators; foundation controls remain Tendril-owned.Wavemap infra operators and the prefix-scoped topology reader.Delivery workflows, runtime secrets, deployment contracts, public docs, or another product’s state.
Legacy recovery backendProtected former Wavemap workload-account backend.No routine writers.Infra operators only during an explicitly declared recovery.Current-state claims, routine preview/update, or application/docs delivery.
Deployment contract storePrivate versioned current-contract objects keyed by Pulumi stack.Approved projection from reviewed Wavemap outputs.Application/docs delivery roles and infra operators through narrow read grants.Raw exports, stack secrets, backend internals, or arbitrary provider output dumps.
Runtime parameter storeAWS SSM Parameter Store under /wavemap/dev/runtime/*.Runtime config population command or approved operator workflow.Runtime host; cloud-plan jobs validate metadata and references without reading secret values.GitHub logs, workflow summaries, rendered env files, or deployment contracts carrying secret values.
GitHub environmentsWorkflow-specific bootstrap values, protected roles, and dispatch controls.Repository operators through GitHub settings.Only the workflow protected by the selected environment.Cloud resource ownership, rendered runtime config, raw state, or long-term evidence.
Private topology artifactsLocal private operator directories and short-retention workflow artifacts.Manual topology workflow or local operator commands.Operators and reviewers.Public docs until a sanitized projection is approved.
Public topology projectionsReviewed docs assets and Mermaid sidecars under the curated docs tree.Human-reviewed docs changes.Anyone reading the docs site.Raw identifiers, workflow evidence, secrets, or private generated candidates.

For deployed dev, this means application and docs delivery should operate from the deployment contract and their own narrow roles. Pulumi credentials, the Pulumi backend passphrase, raw exports, and backend object access remain infra-operator concerns, except for the manual infra-topology ingest role that reads the backend as private evidence.

The deployed-dev model is intentionally cost-conscious and lightweight. A production environment should not inherit that posture without a separate review.

Production planning should revisit at least these hardening items:

AreaProduction Baseline Candidate
Ingress abuse protectionCloudFront Web ACL with rate-based rules, managed rule groups, and alerting sized to expected public traffic.
Account and network isolationDedicated production AWS account, environment-specific trust boundaries, private networking review, and tighter egress.
Pulumi state durabilityKMS-backed encryption review, restore drill, access review, encrypted stack-export backup policy, and lock recovery notes.
Deployment contract handlingSeparate production artifact store, contract publishing writer review, and change evidence tied to release promotion.
Data durabilityManaged database posture, backups, restore testing, monitoring, RTO/RPO ownership, and disaster-recovery expectations.
Release strategyStaging or pre-production gate, cross-environment promotion, rollback proof, and blue/green or canary options.
Observability and auditCloudTrail/log retention, alarms, budget guardrails, workflow evidence retention, and incident-ready dashboards.
Access governanceScheduled review of GitHub environment protections, OIDC subjects, service roles, break-glass credentials, and rotation.
Media and CDN operationsProduction-grade media lifecycle, cache invalidation policy, object retention posture, and CDN monitoring.

Before any workflow performs pulumi up, it must publish a preview summary that a human can review without reading raw Pulumi output first. The summary is a review artifact, not approval by itself.

The first summary format should include:

SectionRequired Content
TargetEnvironment, Pulumi project, stack, cloud provider, selected ref or commit, workflow run, and actor.
Change totalsCounts for creates, updates, replacements, deletes, unchanged resources, and resources that could not be classified.
Review statusno-op, review-required, or blocked, with the reason for the status.
High-risk changesEvery replacement, delete, protected-resource change, data-bearing change, DNS/certificate change, IAM or OIDC trust change, and SSM document change.
Lifecycle classificationResource lifecycle class when known: ephemeral, replaceable, persistent, externally durable, or unclassified.
Output contract changesAdded, removed, renamed, or meaningfully changed stack outputs that deployment, runtime config, smoke, docs hosting, or topology tooling consumes.
Required operator actionApproval needed, runbook to follow, post-apply validation, smoke lane, rollback path, or reason the preview should not be applied.
EvidenceRaw preview artifact location, summary schema version, timestamp, and any sanitization or redaction note.

Use conservative status rules:

  • no-op means the preview has no creates, updates, replacements, or deletes.
  • review-required means the preview is complete enough for human review but still needs explicit approval before apply.
  • blocked means the preview includes an unclassified replacement/delete, a persistent or data-bearing risk without an operator plan, an IAM/trust broadening without permission review, missing raw evidence, or any secret/plaintext leak.

The summary may include Pulumi URNs, resource types, and stable logical names needed for review. Keep raw exports, provider physical IDs, secret values, arbitrary provider inputs/outputs, and workflow logs out of public docs. If a raw artifact is needed for debugging, keep it as private workflow evidence with short retention.

Wavemap is AWS-first, not AWS-only-by-design.

Current convention:

  • Keep provider-neutral vocabulary where the concept is shared: cloudProvider, deploymentEnvironment, runtimeHost, mediaStorage, containerRegistry, ingress, controlPlane, and runtimeConfiguration.
  • Keep provider-specific resources and execution semantics inside provider-owned boundaries.
  • Use AWS-specific names where the code is honestly AWS-specific, such as S3 buckets, ECR repositories, SSM documents, CloudFront distributions, Lambda functions, and Route53 records.
  • Do not add placeholder Azure/GCP directories or workflow paths until a second provider is real enough to clarify the shape.
  • Do not over-abstract shell wrappers that currently execute AWS CLI, Docker, SSM, ECR, S3, CloudFront, or Pulumi behavior.

If a second provider becomes real, prefer a parallel provider adapter selected by typed target/profile logic over one generic wrapper that hides provider differences.

Raw infrastructure evidence can include provider-specific and sensitive-adjacent details. Keep the publication boundary strict:

  • Raw Pulumi exports, DOT labels, workflow evidence, provider physical IDs, and private generated candidates stay private.
  • Sanitized inventory and normalized graph data are review inputs.
  • Public docs receive only reviewed, sanitized projections.
  • Human publication decisions live in infra/topology-processing-reviews/figure-slot-decisions.json.

The topology pipeline is documented in Infra Topology Processing.

  • If a future workflow gains pulumi up, implement the structured preview-summary gate before granting it update authority. The current local launcher already requires a clean full-stack saved plan, exact plan/source digests, and separate human authorization.
  • Add automated checks that approved topology publication paths exist and are referenced from their target docs pages.
  • Model runtime database and volume state as provider-neutral topology nodes.
  • Add scheduled drift checks only when infrastructure churn or operator pain justifies them.
  • Exercise the runtime-host and media-bucket replacement runbooks only when a real preview or planned drill justifies the disruption.