Skip to content

Internal OpenBao Lifecycle ​

This document describes how Hegemony's bundled secret store is started, initialized, and used by the API and worker.

The bundled store is OpenBao, the Linux Foundation fork of HashiCorp Vault. It replaced a bundled Vault, for two reasons: Vault is BUSL-licensed, which is an awkward default to ship with an AGPL product, and OpenBao's static auto-unseal removes the worst part of the old design — an unseal key and a permanent root token sitting in plaintext on a Docker volume.

Secret references did not change. They are still {{ secret('vault://orgs/default/secrets/db/password') }}. The backend keeps the tag vault_internal and the scheme vault, because those are the address every stored reference resolves through; renaming them would break every reference for a cosmetic gain. OpenBao keeps Vault's HTTP API, so the same plugin and the same hvac client serve both.

Overview ​

Available via the openbao-internal compose overlay. (vault-internal still selects it, as a deprecated alias.)

  • Single-node OpenBao for dev and non-HA environments.
  • KV v2 secrets engine mounted at hegemony/.
  • TLS between the API/worker and OpenBao, using a certificate generated at first boot.
  • Auto-unseal from a static seal key, so it comes back ready from every restart.
  • AppRole authentication with two roles, hegemony-api and hegemony-worker, whose credentials are rotated and whose superseded secret IDs are revoked.

When this store is registered as a secret backend, the name you give that backend is a display label following the platform's object name rules. It is independent of the hegemony/ mount path, which is what secret references actually address.

Components ​

ComponentRole
openbao-prepOne-shot. Generates the seal key and the TLS material if absent.
openbaoThe server.
openbao-initLong-lived. Provisions the KV and Transit mounts, policies and AppRoles, then rotates credentials.
deploy/openbao/config.hclServer configuration.
deploy/openbao/policies/api, worker and bootstrap policies.
deploy/openbao/prep.sh, bootstrap.shWhat the two init containers run.

For the compose service definitions and volumes, see deploy/compose/docker-compose.openbao-internal.yml.

Startup sequence ​

  1. openbao-prep runs and exits. It generates a 32-byte auto-unseal key and a CA plus server certificate, each only if not already present, then sets their ownership so OpenBao can read them. An operator who bind-mounts their own key or certificates over those paths keeps them.
  2. OpenBao starts with integrated storage (Raft), a TLS listener on 8200, and the static seal. It auto-unseals, so it reaches a serving state without help.
  3. Healthcheck passes once bao status reports it unsealed. Note that a sealed server now fails the healthcheck — the old Vault healthcheck passed sealedcode=200 and so reported a sealed Vault as healthy.
  4. openbao-init starts as soon as the OpenBao container has started — on service_started, deliberately not service_healthy. Step 3's healthcheck fails while the server is uninitialized, and openbao-init is the only thing that initializes it, so gating on health would deadlock the two on every first boot. bootstrap.sh does its own waiting instead, treating the sealed exit code as "the server answered". It then:
    • On first run, initializes OpenBao, writes the recovery keys to their own volume, and keeps the initial root token in memory only.
    • Enables KV v2 at hegemony/, writes the policies, enables AppRole, and creates the hegemony-api, hegemony-worker and hegemony-bootstrap roles.
    • Mints AppRole credentials and writes them to the shared creds volume.
    • Verifies the scoped hegemony-bootstrap AppRole works, then revokes the root token.
    • Marks itself healthy. The API and worker gate on that.
  5. The API and worker read their credentials from /run/secrets and authenticate over TLS, verifying the generated CA.

That ordered sequence is what docker compose up does. It is not what happens at boot: depends_on conditions are a compose-client feature, and the Docker daemon starts every container carrying a restart policy at once, in no particular order. The stack still converges — see the reboot notes in docs/demo.md — but do not read the order above as a guarantee after a reboot. Auto-unseal is what makes that safe: OpenBao does not need another container to become usable.

Credential rotation ​

openbao-init stays alive after provisioning, and that is its only reason to. On a timer it re-runs the bootstrap, which:

  1. mints a fresh secret ID per role and writes it to the creds volume;
  2. reports itself ready, because everything the API and worker need is now on disk;
  3. waits out a grace period longer than the worker's 60-second backend-config cache, so a process still holding the previous credential has re-read the file;
  4. destroys every superseded secret-ID accessor.

Step 4 is the one the bundled Vault never did — there, every secret ID ever issued stayed valid forever.

Steps 2 and 3 are in that order deliberately. The wait is minutes long and happens on every run except the very first, so announcing readiness after it would mean openbao-init reporting unhealthy on every restart — and the API and worker gate on that health, so docker compose up -d would refuse to start them. The wait is also paid once for the whole run rather than once per role: every credential file is already written when it starts, so one wait covers every reader.

Three settings have to agree, and bootstrap.sh refuses to start if they do not:

VariableDefaultMeaning
OPENBAO_SECRET_ID_TTL24hHow long a secret ID stays valid
OPENBAO_ROTATE_INTERVAL_SECONDS21600 (6h)How often the bootstrap re-runs
OPENBAO_REVOKE_GRACE_SECONDS120Step 3's wait

One rotation cycle is the interval plus the grace, and it has to fit inside the TTL. Shortening the TTL alone gives credentials that expire before their replacements are written, and the symptom — the API and worker failing to authenticate — would surface hours later, nowhere near the setting that caused it. So the check happens at startup instead, and says which of the three to change.

The grace is bounded from below as well: it must be at least 60 seconds, the worker's backend-config cache TTL (_BACKEND_CACHE_TTL_SECONDS in apps/worker/template_resolver.py). A shorter grace destroys the superseded secret ID while a worker that has not yet refreshed that cache is still using it — the mid-flight authentication failure the grace exists to prevent. The floor is a constant rather than a variable, because the value it mirrors is a Python constant too.

The API is not governed by that number, and it is worth being precise about why rather than letting one figure stand for both processes. Its cache (apps/api/services/secret_clients.py) holds backend clients keyed by a config fingerprint and has no TTL at all; it is cleared explicitly, on a backend update or via invalidate_dependent_backend_clients. What keeps a cached API client from pinning a stale credential is not expiry but the *_file indirection described above: the config carries the credential's path, so the path never changes when the file behind it does, and the plugin re-reads it at the next authentication. So 60 seconds is the worker's window specifically — do not read it as a guarantee covering both.

The hegemony-bootstrap credential is the one exception: its secret ID does not expire. That is deliberate. It is openbao-init's only way back in once the root token is revoked, so an expiring one would lock the deployment out of its own provisioning any time the stack stayed down longer than the TTL — recoverable only through the recovery keys. It is re-minted on every run and the superseded accessor destroyed, so a leaked copy still stops working; it just does not expire on a clock. Its policy cannot rewrite its own role or its own policy.

What that credential is worth, precisely. It is not a barrier against reading secrets, and it cannot be made into one. openbao-init issues the api and worker secret IDs and writes them to the shared credentials volume, so it necessarily holds credentials that read the whole hegemony/ tree — before policy enters into it. Its write access to the api role's policy and to the role itself would each be enough on their own, and both are needed for the provisioning it exists to do. Whoever can read the openbao-bootstrap-creds volume is effectively an administrator of the bundled store, and that mount — openbao-init and nothing else — is the control that matters. What the AppRole buys over a stored root token is a credential that is short-lived, rotated, CIDR-bound and revocable without re-initialising the server.

Rotation only reaches a running process because the backend config passes credential file paths rather than resolved values. A {{ file(...) }} placeholder is resolved once, when the API builds the backend client, so a rotated file would never be noticed; the plugin re-reads a *_file path on every authentication.

Storage and volumes ​

One volume per trust boundary, so no container holds more than it needs:

VolumeContentsWho mounts it
openbao-sealThe auto-unseal key, mode 0400openbao-prep writes, openbao reads
openbao-tlsCA and server certificate, plus the server keyopenbao reads; the API and worker mount it read-only for the CA alone
openbao-dataEncrypted integrated storage (Raft)openbao
openbao-credsAPI and worker AppRole credentialsopenbao-init writes, API and worker read-only
openbao-bootstrap-credsThe scoped bootstrap AppRoleopenbao-init only
openbao-recoveryRecovery keys from operator initopenbao-init only

The CA private key stays root-owned and unreadable by OpenBao: OpenBao never needs it, and keeping it out of reach means a compromised server container cannot mint new certificates for itself.

Recovery keys and the seal key ​

Two different things, often confused:

  • The seal key (openbao-seal) decrypts the storage. Losing it loses the data. Back the volume up, or supply your own key by bind-mounting it.
  • The recovery keys (openbao-recovery) are a quorum for OpenBao's root-token ceremony. They do not decrypt storage.

They are no longer sufficient on their own. Since OpenBao 2.6.0 bao operator generate-root calls the authenticated /sys/generate-root-token endpoint rather than the deprecated unauthenticated /sys/generate-root one, so the ceremony needs a token before it can issue one. Run with the recovery keys and nothing else it answers Code: 403 ... permission denied. Treat "the recovery keys get me back in" as false for this deployment until that gap is closed — see Quorum and a snapshot schedule.

bootstrap.sh prints a loud notice when it writes the recovery keys. Move them off the host. For anything beyond a demo, raise OPENBAO_RECOVERY_SHARES and OPENBAO_RECOVERY_THRESHOLD so the shares can be split between people.

Reaching the web UI in development ​

Three separate things keep the bundled OpenBao out of reach, and all three are deliberate:

  1. No UI is served. ui = false in deploy/openbao/config.hcl — the store is driven entirely by the Hegemony API, so the UI would be attack surface and nothing else.
  2. No route from the host. The openbao service declares expose, never ports.
  3. No credential. The root token is revoked at the end of bootstrap, and — see the box above — the recovery keys can no longer mint a replacement.

Dev and demo lift the first two without touching production, and work around the third. Production keeps all three.

Open it ​

bash
task compose:openbao:ui:dev     # or: compose:openbao:ui:demo

That runs openbao-ui, a profiled socat relay that publishes https://127.0.0.1:8200/ui for as long as the command runs; Ctrl-C closes it and removes the container. Nothing is published before or after. It is a TCP passthrough, so OpenBao's own TLS runs end to end and the browser validates the certificate openbao-prep generated — self-signed, so expect a warning.

The UI itself is enabled by BAO_UI=true in docker-compose.openbao-internal-dev.yml and the demo-data repository's docker-compose.demo.yml rather than by editing config.hcl, which is what keeps production unaffected. BAO_UI is OpenBao's own documented override for the ui setting. dc.sh and the compose Taskfile add the dev file right after the openbao-internal overlay whenever ENV=dev, and never in production; dc.sh does the same when the overlay is named through EXTRA_FILES. If you run docker compose yourself, list it after the overlay to get the UI; without it the dev stack still starts, with the UI off.

Get a token ​

bash
task compose:openbao:token:dev   # or: compose:openbao:token:demo

This logs in with the hegemony-api AppRole — the same credential the Hegemony API uses — and prints the resulting token. Paste it into the UI's token login.

It is deliberately not a root token, and not for want of trying: there is no root token to hand out (bootstrap revokes it) and the recovery keys cannot mint one, for the reason in the box above. The API role is what exists.

That turns out to be the right scope anyway. Its policy covers Hegemony's own KV tree — browse, read, write — which is what anyone opening this UI came to look at. Everything else, including OpenBao's administrative sections, returns permission denied.

The token carries the role's token_ttl (20 minutes by default, set with OPENBAO_TOKEN_TTL), so an abandoned one expires on its own. To end it sooner:

bash
BAO_TOKEN=<token> bao token revoke -self

The AppRole is bound to the OpenBao network's subnet, for both the login and the token's own use. That still works from a browser on the host because the requests arrive through the openbao-ui relay, which is a TCP passthrough on that network, so OpenBao sees the relay's address rather than yours.

In production ​

Both services are profiled, so up never starts either one, in any environment. The relay would work if an operator ran it, exactly as temporal-ui does; the UI would not render, because ui = false stands there. Reaching for it should be a deliberate act, which is why neither is wired into any service list.

If you only need to read a secret rather than browse, task compose:demo:exec into a container that already holds an AppRole and use the CLI — no relay and no root ceremony.

The demo's external Vault is a different service and is reachable without any of this, on purpose: http://localhost:8201/ui with token demo-external-root. It runs server -dev with in-memory storage and holds only synthetic demo secrets, so it is the one to poke at when you just want to see a secret backend's UI.

Reset behavior ​

  • Wiping the volumes resets everything, including the seal key: the data is gone, not merely locked.
  • After a reset, openbao-prep generates a new seal key and certificate and openbao-init re-provisions from scratch.
  • Restart or recreate api and worker once openbao-init is healthy again, so they pick up the new credentials rather than stale ones.

Migrating from a previous Vault-backed install ​

There is no in-place data migration. OpenBao's documented migration covers Vault CE 1.14.x on Raft storage; the bundled Vault was 1.19 on file storage, which is outside it. A new OpenBao starts empty.

The backend row migrates itself — the API reconciles its type, name and config on every startup, so no database migration is needed — but the secret values have to be moved:

  1. Before upgrading, with the old stack still running, export from Vault:

    bash
    docker compose ... exec vault sh -c '
      vault kv list -format=json hegemony/metadata/orgs' # walk the tree

    Read each path with vault kv get -format=json hegemony/<path> and keep the output somewhere safe. Treat that export as live credential material.

  2. Bring up the new stack and let openbao-init finish.

  3. Write each secret back, either through the Hegemony API (preferred — it keeps the metadata rows in step) or with bao kv put hegemony/<path> k=v.

  4. Run a flow that resolves a vault:// reference to confirm.

Wipe the old Vault volumes only once you have confirmed the new store works.

Transit and managed state keys ​

openbao-init also enables the Transit engine at transit/. The API uses it to protect managed Terraform/OpenTofu state: each stored version gets a fresh data key from transit/datakey/plaintext/<key>, and only the wrapped key is kept in PostgreSQL. There is one key per organization, hegemony-org-<org id>-tf-state, created by the API on first use. The api policy allows creating keys on this mount and using only those keys, and nothing below a key: no configuration (so no deletion or export), rotation or trimming. It cannot read a key's metadata either. It names the key path with transit/keys/+, which matches one path segment, because creating a key needs update and a * glob there would also match the paths below every key.

Back up OpenBao and PostgreSQL together: a state version can only be read with the key that wrapped it. The bundled backup task covers neither openbao-data nor openbao-seal; see "Encryption and Keys" in Managed Terraform/OpenTofu state for what to copy.

An OpenBao set up before this engine was added keeps its older bootstrap policy, because that policy changes only while the initial root token exists. There openbao-init logs that it could not enable Transit, with OpenBao's reason, and managed state gets 503 until you choose the platform key: set HEGEMONY_TF_STATE_KEY_BACKEND=none and HEGEMONY_TF_STATE_ENCRYPTION_KEY. Enabling Transit afterwards needs a token with sys/mounts write, which such a deployment no longer has on OpenBao 2.6 or later (see Quorum and a snapshot schedule for why an operator token cannot be minted there).

Operational notes ​

  • A sealed OpenBao does not make the API unhealthy: the API resolves secrets lazily, so a store that failed to come up presents as flow steps failing on secrets, not as a down stack. Check openbao's health first, then openbao-init's logs.
  • depends_on gates container starts only. An API or worker that is already running is not stopped by openbao-init going unhealthy; it keeps running, and its secret resolution fails until the problem clears.
  • AppRole token TTLs are deliberately short (20 minutes, 1 hour maximum) and the API and worker re-authenticate as needed.
  • The bundled store is meant for development, evaluation and explicit opt-in production use. External backends can be configured through the Settings UI, whose Secret Backends pages every signed-in user can open; the add, edit, test and delete controls are enabled only for admins.
  • The backend's page in the Settings UI carries a History card with the most recent audit log entries about the backend row: who registered it, who changed its address or credentials, and who removed it. The card is shown only to callers who may read the audit log.

Troubleshooting quick checks ​

  • OpenBao will not start. Read its logs first. The three most likely causes are all pre-flight issues, each with a one-line fix:
    • cannot read the seal key — OPENBAO_UID/OPENBAO_GID do not match the uid the image runs as. The defaults (100:1000) match openbao/openbao:2.6.2; re-check after an image bump with docker run --rm --entrypoint sh openbao/openbao:<tag> -c 'id openbao' and set them in the environment if they differ.
    • OpenBao has dropped support for mlock — a disable_mlock line has been added back to deploy/openbao/config.hcl. Remove it: OpenBao refuses to start with the setting present at all, in either direction. See Memory locking below for what to do instead.
    • unknown seal type "static" — the pinned image predates the static seal. Move to 2.5 or newer.
  • openbao-init stalls. Check OpenBao is reachable over TLS and that the CA in /openbao/tls is the one the certificate was signed by.
  • AppRole auth failures. docker compose ... restart openbao-init re-runs the bootstrap and re-mints credentials. If you reset the volumes, restart api and worker once it is healthy again.
  • Everything is healthy but secrets fail after a network change. The AppRole credentials are bound to the OpenBao network's CIDR. If you changed OPENBAO_NETWORK_SUBNET, change OPENBAO_BOUND_CIDRS to match and restart openbao-init.

What is already hardened, and what is left ​

The bundled Vault this replaced needed a long list of changes before production. Most of them are now the default:

ConcernStatus
TLS in transitDone — generated certificate, verified by the API and worker
Unseal key handlingDone — auto-unseal; no unseal key on a shared volume
Root tokenDone — revoked after bootstrap; a scoped AppRole takes over
Credential lifetimeDone — short token TTLs, expiring secret IDs, superseded ones revoked
Network exposureDone — no published port; credentials bound to a dedicated network
Audit loggingDone — file audit device to stdout, from boot
Storage engineDone — integrated storage (Raft); snapshots available
Memory lockingNot available — OpenBao dropped mlock; see Memory locking
Container privilegesDone — every capability dropped, no new privileges, read-only root filesystem

What remains is genuinely site-specific:

Memory locking ​

This one is a step backwards from the Vault it replaces, and it cannot be fixed in this configuration.

OpenBao has removed mlock support. disable_mlock is not merely ignored — the server refuses to start if the setting appears at all:

text
error loading configuration: OpenBao has dropped support for mlock.
Please remove the line "disable_mlock" = false from your config and disable
or encrypt swap instead.

So the root key can be paged out to swap, and the protection has to come from the host. Do one of these on any machine running the bundled backend:

bash
sudo swapoff -a          # then remove the swap entry from /etc/fstab

or encrypt the swap device (dm-crypt with a random key per boot, which most distributions offer as a swap entry in /etc/crypttab).

The container drops IPC_LOCK along with every other capability, since with mlock gone it would grant a privilege nothing uses.

1) A real certificate ​

The generated certificate is self-signed. It encrypts the hop and exercises the verification path; it is not identity. Replace server.crt/server.key in the openbao-tls volume with certificates from your own CA, and point the backend's ca_cert_file at that CA.

2) A seal you do not keep next to the data ​

Static-key auto-unseal puts the seal key on the same host as the storage, so host root gets both. That is inherent to a self-contained bundled store. For production, move to a transit seal (backed by another OpenBao) or a cloud KMS, and delete the static key.

3) Quorum and a snapshot schedule ​

The bundled backend already runs on integrated storage (Raft) rather than the file backend, which OpenBao deprecated and removes in 2.7.0. What it does not have is quorum: it is one node, so a lost volume is lost data.

Two things to add for production:

Snapshots. Raft has a snapshot command; the file backend had none.

⚠️ The ceremony below does not work on OpenBao 2.6.0 or later, and this runbook has no replacement yet. bao operator generate-root now calls the authenticated /sys/generate-root-token endpoint, so step 1 fails with Code: 403 ... permission denied when the recovery keys are all you hold. The steps are kept because the shape of the ceremony is unchanged and the rest of the runbook — the snapshot, the transfer, the cleanup chaining — is still correct once you have an operator token by some other means. Do not plan an outage around this section until that gap is closed.

Taking one needs an operator credential, and the bundled setup deliberately leaves none lying around: the root token is revoked at the end of bootstrap, and the hegemony-bootstrap AppRole has no sys/storage/raft/snapshot access — it provisions mounts and roles, and giving it the ability to read the whole store would defeat the point of scoping it. So the first step is minting an operator token from the recovery keys, which is what they are for.

Everything below runs through the repo's compose wrapper, so set the stack up once:

bash
cd deploy/compose
export ENV=prod SERVICES=auth,s3,otel,openbao-internal   # whatever stack you run

Mint the operator token from the recovery keys. -T is what lets a value be captured into a shell variable; the key-submission step deliberately runs with a TTY so each holder types their key at a hidden prompt instead of leaving it in ps output and shell history:

bash
# 1. An OTP only you hold. It is what decodes the token at the end.
OTP=$(./dc.sh exec -T openbao bao operator generate-root -generate-otp | tr -d '\r\n')

# 2. Start the ceremony. Note the "Nonce" it prints and set it below.
./dc.sh exec -T openbao bao operator generate-root -init -otp="$OTP"
NONCE=<the nonce printed above>

# 3. Each recovery-key holder submits their key at the prompt. The keys were written
#    to the openbao-recovery volume at init; at the dev/demo threshold of 1 this runs
#    once. The final submission prints "Encoded Token" -- set it below.
./dc.sh exec openbao bao operator generate-root -nonce="$NONCE"
ENCODED=<the encoded token printed above>

# 4. Decode it. This token exists only in this shell; nothing persists it.
BAO_TOKEN=$(./dc.sh exec -T openbao bao operator generate-root \
    -decode="$ENCODED" -otp="$OTP" | tr -d '\r\n')

Setting BAO_TOKEN on the host does not put it in the container's environment, so it has to be handed to each exec explicitly. Passing it as -e BAO_TOKEN=... would expose it in ps on the host, so send it on stdin instead:

bash
# Runs one command in the openbao container as the operator.
bao_op() {
    printf '%s\n' "$BAO_TOKEN" | ./dc.sh exec -T openbao sh -c '
        IFS= read -r BAO_TOKEN
        export BAO_TOKEN
        exec "$@"' _ "$@"
}

Then take the snapshot. Each cleanup is chained behind the transfer before it, so a failed copy never deletes the only readable snapshot:

bash
SNAPSHOT="/tmp/bao-$(date +%F-%H%M).snap"

bao_op bao operator raft snapshot save /tmp/bao.snap
./dc.sh cp openbao:/tmp/bao.snap "$SNAPSHOT" \
    && bao_op rm -f /tmp/bao.snap          # leave nothing in the container

# The host is not a backup: it holds the Raft volume too, so a host loss takes both.
# Substitute your own upload -- it must exit non-zero when the transfer fails, or the
# `&&` below will delete the host copy of a snapshot that never arrived.
aws s3 cp "$SNAPSHOT" s3://your-backup-bucket/openbao/ \
    && rm -f "$SNAPSHOT"

Revoke the operator token when the run is done, so the store goes back to having no standing admin credential:

bash
bao_op bao token revoke -self
unset BAO_TOKEN

For an unattended schedule, prefer a dedicated AppRole whose policy grants exactly sys/storage/raft/snapshot (read for save, update for restore) over keeping a generated root token alive.

A snapshot restores with bao operator raft snapshot restore. It is encrypted with the same seal key, so a snapshot is not a substitute for backing up openbao-seal — without that key the snapshot is unreadable.

More nodes. Three, not two: quorum is a majority, so two nodes tolerate no failures and add only a way to lose one. Each node needs its own node_id (or BAO_RAFT_NODE_ID), an api_addr/cluster_addr that other nodes can actually route to, and retry_join blocks naming its peers.

4) Ship the audit log ​

The audit device writes to stdout, which is right for a single container but is not retention. Ship container logs to your central logging, restrict who can read them, and match your retention policy.

5) Tighten AppRole issuance further ​

secret_id_num_uses is unlimited, because single-use secret IDs need an agent alongside each consumer to keep re-issuing them. If you run such an agent, set secret_id_num_uses=1 and shorten secret_id_ttl. Consider response-wrapped secret ID delivery.

6) Operational security ​

Run on a host with swap disabled or encrypted and core dumps off, restrict inbound access to the API and worker, and keep OpenBao patched.

Released as open source under the AGPL-3.0-or-later license. Development is sponsored by Rexonix s.r.o.. Contact — [email protected].