The fastest way to try the reference topology on a laptop.
make deploy # 3-node Vault cluster, auto-unsealed
make status # check cluster/unseal status
make destroy # tear it downThe cluster serves TLS. bootstrap-dev-cluster.sh calls
scripts/generate-dev-certs.sh first, which issues a local CA and a
certificate per node into docker/dev/tls/ — gitignored, and regenerated
with --force if they ever need replacing. Clients need the CA:
export VAULT_ADDR=https://127.0.0.1:8200
export VAULT_CACERT=$PWD/docker/dev/tls/ca.crtmake deploy runs scripts/bootstrap-dev-cluster.sh, which brings up a
standalone Vault instance (vault-unseal) as a Transit auto-unseal
backend, then starts and initializes the 3-node cluster against it — the
same seal stanza shape the AWS/Azure profiles below use, just pointed at
something that needs no cloud account. vault-unseal itself is still
unsealed the manual, Shamir way — something has to be the root of trust.
See docs/auto-unseal.md for the full picture.
Neither has ever been applied to a real account. The AWS profile is applied and destroyed against an emulated AWS API on every PR, which settles that it applies at all; nothing in that run boots, so it says nothing about the cluster. Run the pre-flight first — it checks credentials, the inputs that fail late, quota and cost, and applies nothing:
./scripts/preflight-cloud.sh --cloud awsThen read cloud-apply.md, which lists what to verify
while the cluster is up and how to tear it down afterwards.
terraform destroy alone does not fully work on either profile.
cd terraform/aws
terraform init
terraform plan -out=plan.tfplan
terraform apply plan.tfplanProvisions:
- A VPC with public and private subnets across
az_countavailability zones, one NAT gateway per AZ, and an S3 gateway endpoint - An autoscaling group of Vault nodes in the private subnets, sized to
node_count - A network load balancer on port 8200, internal by default
- A KMS key for auto-unseal, plus the instance role that uses it
- An S3 bucket for Raft snapshots, versioned and lifecycle-expired
docs/security.md commits to TLS terminating at the Vault
process rather than being offloaded. An application load balancer can't
do that — it terminates the client's TLS and opens a separate connection
to the backend, so plaintext exists inside the load balancer. An NLB
forwards TCP untouched, so the client's TLS session runs end to end with
Vault and the load balancer never holds a certificate or sees a token.
The health check still speaks HTTPS to /v1/sys/health, accepting both
200 (active) and 429 (standby), so every unsealed node stays in the
pool and writes get forwarded to the leader.
The user-data writes a Vault config with a TLS listener but does not
issue certificates — how you get them is deployment-specific (an internal
CA, ACM Private CA, or Vault's own PKI engine once a first cluster
exists). Until they are in place at /etc/vault.d/tls/, Vault will not
start. That is deliberate: a Vault serving plaintext is worse than one
that refuses to boot.
Delivering them is what the Ansible layer is for; see Handing off to Ansible below.
Three NAT gateways at roughly $32/month each are the bulk of the idle
cost. Dropping az_count to 2, or sharing a single NAT, trades that
against AZ independence.
cd terraform/azure
terraform init
terraform plan -out=plan.tfplan \
-var="ssh_public_key=$(cat ~/.ssh/id_ed25519.pub)"
terraform apply plan.tfplanMirrors the AWS layout: a VNet with separate node and load balancer
subnets, a VM scale set sized to node_count, a Standard load balancer
on 8200, Key Vault auto-unseal, and a storage account for Raft snapshots.
ssh_public_key is required — Azure will not create a Linux scale set
with neither a password nor a key.
Same group_vars/vault_nodes_azure.yml.example step as AWS before
running the playbook.
The two are meant to behave the same, but the mechanisms differ in ways worth knowing:
- Subnets are regional, not zonal. One subnet spans the region and zone spread is a property of the scale set, so there is a single node subnet rather than one per zone.
- The health probe has no status-code matcher. Azure probes accept
200-299 and nothing else, while Vault answers 429 on a standby. The
probe passes
standbyok=trueso Vault answers 200 for a healthy standby instead — without it Azure ejects every standby and only the leader serves traffic. - Outbound needs an explicit NAT gateway. Azure's default outbound access is being retired, and relying on it means nodes lose internet access on a date outside your control.
- Names are length-limited and globally unique. Key Vault and storage
account names are capped at 24 characters across all of Azure, so both
are truncated and given a random suffix rather than derived from
cluster_namealone.
The NAT gateway and the Standard load balancer are the bulk of the idle
cost, in the same range as the AWS profile's NAT gateways. Premium OS
disks add to it; os_disk_size_gb and vm_size are the levers.
Terraform builds hosts that cannot serve until something gives them their
certificates and their configuration. That something is the Ansible
layer, and the two halves have to agree on a dozen values — the KMS key
id, the subscription id, the scale set name, the region. Copying them by
hand out of terraform output works exactly once.
./scripts/terraform-to-ansible.sh --cloud awsThat reads terraform output -json and writes
ansible/group_vars/vault_nodes.yml. Re-run it after any apply rather
than editing the file — a hand edit drifts from the infrastructure it
describes and nothing catches that. It refuses to overwrite an existing
file unless you pass --force, and it aborts without writing anything if
an output it needs is missing, rather than emitting a null that becomes
a Vault which starts and cannot unseal.
Then run the playbook:
cd ansible && ansible-playbook -i inventory/aws.yml playbooks/site.ymlSubstitute inventory/azure.yml for the Azure profile.
inventory/aws.yml and inventory/azure.yml discover nodes through the
cloud API by tag, not from a list of addresses. The autoscaling group and
the scale set both replace instances, so a static inventory is wrong the
first time a node is recycled — and wrong silently: the playbook
succeeds against hosts that no longer exist and never touches the ones
that do.
On AWS the tag they filter on is the same one Raft's auto_join uses, so
cluster formation and configuration management break together rather than
one drifting away from the other.
On Azure they are independent. go-discover's Azure provider rejects a
mix of tag and scale-set selectors, so retry_join enumerates the scale
set by resource group and name and never looks at tags. An empty
inventory there says nothing about whether the cluster formed, and a
healthy cluster is no evidence the inventory works.
Nodes sit in private subnets with no public address, so reaching them needs SSM, a bastion, or a VPN.
The role expects to find certificates on the control machine and copies them to each node:
ansible/files/tls/ca.crt
ansible/files/tls/<inventory_hostname>.crt
ansible/files/tls/<inventory_hostname>.key
Per-node leaves rather than one shared certificate, matching what
scripts/generate-dev-certs.sh produces locally. Override
vault_tls_ca_src, vault_tls_cert_src, and vault_tls_key_src to
point elsewhere. The role verifies each certificate actually matches the
host it lands on, because the alternative failure surfaces later as a
Raft join error that reads like a network problem.
The autoscaling group ships with health_check_type = "EC2", and that is
a deliberate compromise you are expected to undo.
EC2 health only asks whether the instance is running. That is the only
thing the group can usefully ask before this step: Terraform does not
issue certificates, so until the playbook has run, Vault does not start,
the load balancer's check cannot pass, and an ELB health check would
mark every instance unhealthy at the end of its grace period, terminate
it, and launch a replacement that does the same. A bare apply would never
converge, and would bill for the privilege.
Once the playbook has converged and the nodes are serving:
terraform apply -var health_check_type=ELBNow a node that is running but sealed, wedged, or out of the Raft
quorum is replaced, which EC2 health cannot see. Leaving it on EC2
means an instance can sit up and useless indefinitely.
Confirm the group agrees before relying on it:
Q='AutoScalingGroups[].[AutoScalingGroupName,HealthCheckType]'
aws autoscaling describe-auto-scaling-groups --query "$Q" --output textBoth settings are asserted in terraform/aws/tests/cluster.tftest.hcl --
that the default is EC2, and that ELB is reachable -- so neither half
can be dropped without a test failing.
tests/ansible/run-tests.sh exercises the handoff against saved
terraform output -json fixtures: the generated group_vars, the
rendered vault.hcl for both clouds, and the case where no cloud is
configured. It needs no credentials and runs in CI.
It does not prove the playbook converges against real hosts. Neither cloud profile has been applied to a real account, and the emulated apply covers Terraform only — it never reaches the Ansible layer. See Provider lock files and the note in the README.
terraform/aws and terraform/azure each commit a
.terraform.lock.hcl. It pins the exact provider versions and records
their checksums, which does two things: an upstream provider release
can't change what CI builds, and a substituted or tampered provider
can't install silently.
The lock records checksums per platform, and Terraform refuses to
run on a platform the lock doesn't cover. These are locked for
linux_amd64 (CI), darwin_arm64, and windows_amd64. Working on
something else — an Intel Mac, an ARM Linux runner — means adding it:
cd terraform/aws
terraform providers lock \
-platform=linux_amd64 \
-platform=darwin_arm64 \
-platform=windows_amd64 \
-platform=linux_arm64 # the one being addedList every platform to keep, not just the new one: the command replaces the set rather than adding to it.
To take a newer provider version, widen the constraint in the module's
required_providers block and re-run the same command. Don't hand-edit
the file — the checksums are the point of it.
Note that terraform init alone writes a lock for only the current
platform, which is why the command above exists. Committing an
init-generated lock is the usual way this gets broken: it works locally
and then fails everywhere else.
- Initialize Vault (
vault operator init) — do this exactly once per cluster, and distribute unseal/recovery keys per your organization's policy. - Apply baseline policies from
examples/policies/. - Enable and configure the audit device.
- Confirm Raft peer status:
vault operator raft list-peers.