Managed Terraform and OpenTofu State
Hegemony can store Terraform and OpenTofu state for you. tf.plan and tf.apply steps use it by default (State: managed), and people can use the same states from their own machines. A state is:
- named and belongs to one organization (for example
network-core); - locked while a plan or apply runs, so two runs never write it at once, and optionally held for a run from its plan until its apply;
- versioned, so you can see every change, download a version, or roll back;
- importable, so a configuration whose state lives elsewhere can move in;
- encrypted with a key kept outside the database.
It is served over the tools' standard http backend, so the configuration needs no backend block and no cloud account.
In a Flow
Give the tf.plan step a State name. The step replaces any backend block in the configuration with the managed backend and gets a short-lived token for that one state, which ends with the step. The token lasts as long as the step's timeout plus a margin, but at most 24 hours: a step that runs longer cannot save its state after that. The tf.apply step that applies the plan opens the same state again. The saved plan waits in the run's shared workspace until the run ends, however long an approval between the two steps takes. Every flow using the same name shares one state and its lock.
Keep a flow that applies changes to one run at a time (the flow's run limits), so a second run does not queue behind the lock. When a step gives up waiting for a lock, the tool's "Error acquiring the state lock" message shows the lock's Info, which names the run and step (or the user) that holds it and since when.
Holding the State from Plan to Apply
The lock covers the plan and the apply themselves, not the time between them. While a run waits for an approval, another run or a person can change the state. The apply then refuses the saved plan as out of date ("Saved plan is stale"), changes nothing, and the run has to plan again.
To keep the state for one run from its plan to its apply, turn on Hold the state until apply in the tf.plan step (managed state only). The plan's lock then also holds the state for the run, in the same step, so nothing can come in between:
- Only that run's steps can lock or write the state. Other runs, people using a token, roll back and delete are refused as if the state were locked; a tool waiting for the lock prints the hold's
Info, which names the run. Reading the state still works. - The hold ends when the run's apply releases its lock, when the run ends (however it ends, including a rejected approval), or when someone force-unlocks the state. A run whose plan finds no changes and skips its apply holds the state until it ends.
A hold blocks every other flow that uses the state for as long as the approval waits, so use it where a plan must be applied as approved, and keep such approvals short. The state list and the state's page show the hold and link to the run.
A lock taken by a run step is released when that step's container is gone, however the step ends, so a container that died mid-apply never leaves the state stuck. The lock belongs to the step's token: a retried step gets a new token, so the cleanup of an earlier attempt never frees a later attempt's lock. Cancelling a run does not release it at once: the worker stops a running step only at its next heartbeat, and until then the step can still change infrastructure. So the step keeps its lock, and can still save its state, until the worker reports its container gone, or at most HEGEMONY_TF_STATE_RUN_END_GRACE_SECONDS (600 by default) after the run ended. Once the step's token no longer works, the state list shows the state as unlocked and the next lock, write, roll back or delete takes the lock over; releasing it is not a force unlock. Stopping a step mid-apply ends Terraform or OpenTofu abruptly: a resource being changed at that moment may be missing from the state and need import.
If the worker itself dies, nothing reports the container gone. Container steps have a two-minute heartbeat timeout, so the step fails (or is retried) about two minutes after the crash, and the run can end; the lock is held until the run ends plus the grace. On a sandbox that outlives the worker, the container keeps running, so two things stop it:
- Every worker sweeps its sandbox every minute (
HEGEMONY_CONTAINER_SWEEP_INTERVAL_SECONDS). A step container whose run ended at least two minutes ago gets SIGTERM, then up to its stop timeout (at most two minutes), and is removed. That is about five minutes after the run's end at most, inside the grace (600 seconds by default), while the step's state access still works. The run script passes the SIGTERM on, and the stop timeout is the step's Stop grace period, so the tool uses it to save the state of the changes it made and unlock. - The
tf.planandtf.applyrun scripts stop the tool themselves before the step's timeout: SIGTERM a Stop grace period before it (at most half the timeout), on which the tool starts no new changes and saves the state of those done, then SIGKILL once that grace has passed, and not before five minutes. This covers a sandbox no worker sweeps, and needs atimeoutcommand in the image (the default images have one).
If you lower HEGEMONY_TF_STATE_RUN_END_GRACE_SECONDS, keep it longer than the sweep can take: two minutes, plus one sweep interval, plus the stop timeout (at most two minutes), plus time for the API and Docker calls. With the defaults, keep it at 480 or more; a shorter grace can free the lock while the tool still runs. With the sweep turned off, the tool can run until the step's timeout, long after the grace; then, after a worker crash, stop that run's step containers on the sandbox (they carry the label hegemony.run_id=<run id>) before running the flow again.
When an apply changes resources but the tool cannot save the resulting state, tf.apply saves it again from the tool's errored.tfstate; if that fails too, it keeps the state as the run artifact errored.tfstate.gz (see the tf.apply step's documentation).
From Your Machine
Use a personal access token with the operator role (or higher) as the password; any username works. Put the backend settings in a file that is not committed, for example backend.hcl:
address = "https://hegemony.example.com/api/tf-states/network-core/state"
lock_address = "https://hegemony.example.com/api/tf-states/network-core/lock"
unlock_address = "https://hegemony.example.com/api/tf-states/network-core/lock"
lock_method = "POST"
unlock_method = "DELETE"
username = "me"Then pass the token in the environment and initialize:
export TF_HTTP_PASSWORD="$HEGEMONY_TOKEN"
tofu init -backend-config=backend.hclWith a terraform { backend "http" {} } block in the configuration, tofu init -migrate-state moves an existing state into Hegemony. The token decides the organization: a token bound to an organization uses that one. You can also upload the state file in the web interface (see Import an Existing State).
The web interface's address forwards /api/… to the API through its web server. The shipped web image accepts requests up to 128 MiB under /api/tf-states, the most HEGEMONY_TF_STATE_MAX_BYTES allows, and keeps nginx's 1 MiB for every other API route. A reverse proxy of your own in front of Hegemony needs a limit at least as high as that setting for that path (nginx accepts 1 MiB unless configured otherwise). In nginx, match the path with a regex, as the web image does (location ~ ^/api/tf-states(/|$)): a prefix location ending in / makes nginx redirect /api/tf-states, the state list, to an absolute URL without your port.
Import an Existing State
To move a configuration whose state lives elsewhere (a local file, an S3 bucket, Terraform Cloud) into Hegemony, import its state file:
- Stop everything that writes the old state, then save it:
terraform state pull > network-core.tfstate(ortofu state pull). - In Artifacts → Terraform States, choose Import state, pick the file, and name the new state. The name starts as the file's name; change it if you like.
- Give that name to the
tf.planstep as its State (or put it inbackend.hcl), and stop using the old backend, so the two copies do not drift apart.
The file becomes the state's first version, with source import, and keeps its serial and lineage, so the next plan continues from it. An import only creates a state: a name already in use is refused, so an import never changes an existing state. To replace one, delete it first, or write a new version with tofu state push, which the checks on every write apply to. The file must be a Terraform or OpenTofu state document (an OpenTofu-encrypted state is stored as it is) of at most HEGEMONY_TF_STATE_MAX_BYTES. Importing needs the operator role, and each import is in the audit log.
To import many states, a script can call the API with a personal access token (as a Bearer token; Basic works only on the backend paths):
curl -X POST "https://hegemony.example.com/api/tf-states/network-core/import" \
-H "Authorization: Bearer $HEGEMONY_TOKEN" \
-H "Content-Type: application/json" \
--data-binary @network-core.tfstateBrowse, Roll Back, Unlock
Artifacts → Terraform States in the sidebar lists the organization's states. Each row shows the latest version, serial, size, when it last changed, and who holds the lock, or which run holds the state until it applies. A lock or hold of a run links to that run. Every state is stored encrypted (see Encryption and Keys); states that OpenTofu also encrypts itself are marked OpenTofu-encrypted.
Open a state to see its lineage, its lock, and its versions, with who wrote each one. From there you can download a version, roll back to an older one, force-unlock the state, or delete it. You only see the actions your role allows (see the table below). Roll back and delete are not available while the state is locked or held for a run. Roll back, force unlock and delete ask you to confirm first. Force unlock also ends a run's hold; that run's apply then refuses its plan if the state changed meanwhile.
The page uses these API routes, which never return state content in a listing:
| Route | What it does | Role |
|---|---|---|
GET /tf-states | States, their latest version, size, current lock, and a run's hold | viewer |
GET /tf-states/{name}/versions | Versions: serial, size, SHA-256, who wrote it (run and step, or user) | viewer |
GET /tf-states/{name}/versions/{version}/content | Download one version | operator |
POST /tf-states/{name}/import | Create a state from a state file (a name not in use only) | operator |
POST /tf-states/{name}/rollback | Make an older version current again, as a new version | operator |
POST /tf-states/{name}/force-unlock | Release a lock held by someone else, and end a run's hold | admin |
DELETE /tf-states/{name} | Delete the state and all its versions (only when not locked or held) | admin |
A state can hold secrets (passwords and keys of the resources it describes), so reading or downloading one needs the operator role. tofu force-unlock and terraform force-unlock also work, for whoever may use Force unlock (the tf_state:manage permission, admin unless an administrator allocated it otherwise; see Permissions). Only the holder of a lock (the same user, or the same run step token) writes under it or releases it without that permission: releasing someone else's lock is a force unlock even with its lock ID, which the state list and every lock conflict show. A run step never forces. An unlock that names a lock ID must name the current lock, so an unlock meant for an old lock never frees the one that replaced it; the Force unlock action and an empty unlock request release whatever lock is held, and a run's hold. A hold has an ID of its own, hold-<run id>: tofu force-unlock hold-<run id> ends the hold and leaves a lock alone. Writes of new versions, imports, rollbacks, force unlocks, deletions and downloads of a version are in the audit log.
A tf.plan or tf.apply step's token acts with the operator role. Allocating tf_state:read or tf_state:write above operator, or disabling them, also refuses those steps.
Writes made under one lock update a single version: an apply saves the state every few seconds, and the history keeps one version per operation. That lasts until the lock is released: a lock taken later, even with the same lock ID, starts a new version, so a released version never changes. The newest HEGEMONY_TF_STATE_KEEP_VERSIONS versions are kept (50 by default).
Encryption and Keys
Each version is compressed and encrypted with AES-256-GCM under its own data key. Only the data key's encrypted ("wrapped") form is stored next to it, so a database dump alone reveals nothing. The key that unwraps it comes from, in order:
- the secret backend named in
HEGEMONY_TF_STATE_KEY_BACKEND(its tag), through its Transit engine (OpenBao or HashiCorp Vault); - otherwise the internal OpenBao's Transit engine, with one key per organization;
- otherwise the platform key
HEGEMONY_TF_STATE_ENCRYPTION_KEY: 32 random bytes, base64-encoded (openssl rand -base64 32).
The choice follows this configuration only. When Transit fails (OpenBao is sealed, unreachable, or refuses the request), writes get 503 and the tools retry; Hegemony never switches to another key on its own, because the key a version is sealed with decides what you must back up to read it. Set HEGEMONY_TF_STATE_KEY_BACKEND=none to always use the platform key, for example with an internal OpenBao set up before Transit support, which cannot enable the engine itself (see the OpenBao page). With none of these, managed state is unavailable and tf.plan fails with that reason before it starts.
Each version records which key wrapped it (the version list shows transit or platform, and so does the audit entry of each write), so changing the setting later keeps older versions readable as long as their key still exists.
Back up the key together with the database. A lost key means lost states.
- Platform key: keep it wherever you keep the database backup's secrets. Never change it; versions it wrapped become unreadable.
- Internal OpenBao: the Transit keys live in OpenBao's storage (the
openbao-datavolume), which only its seal key (theopenbao-sealvolume) can decrypt. The bundled backup (ACTION=backup task compose:objectstore) does not include either, and a Raft snapshot needs an operator token that cannot be obtained on OpenBao 2.6 or later (see the OpenBao page). Back both volumes up yourself in the same window as the database, with OpenBao stopped, for exampledocker run --rm -v <project>_openbao-data:/data:ro -v <project>_openbao-seal:/seal:ro -v "$PWD":/backup alpine tar czf /backup/openbao.tgz /data /seal(managed state and secret reads fail while it is stopped). If you cannot, use the platform key instead. - External Vault or OpenBao: follow its own backup procedure, and keep its backups as long as the database backups that need them.
For state that Hegemony never sees in clear, add OpenTofu's own encryption with the step's State encryption key: OpenTofu then encrypts the state before sending it, and Hegemony stores the encrypted document as it is (it cannot check its serial or lineage then).
Checks on Every Write
- A locked state accepts writes only from the lock's holder. A write with an ID must name the current lock; the holder may leave the ID out (a save outside the locked operation, such as
tf.applysaving a state the tool failed to save), which becomes a version of its own. - A state from another configuration (a different lineage) is refused.
- An older serial than the stored one is refused (a stale state).
- The state must be at most
HEGEMONY_TF_STATE_MAX_BYTES(32 MiB by default, 128 MiB at most). A larger upload is refused as it arrives, with or without aContent-Length.
A lock request is checked too. Its lock document must be a JSON object of at most 16 KiB (413 when larger) with an ID of 1-255 characters, and every field the tools send (ID, Operation, Info, Who, Version, Created, Path) must be a string (400 otherwise). Other fields are dropped, and so is a Created that is not an RFC 3339 time, since other lockers read it back as one.
The API holds a state in memory while it checks and encrypts it: about 2.5 times the state's size (up to about 5 times when the state holds non-Latin text), for every state being saved at the same moment. With the 32 MiB limit that is under 100 MiB per save; with 128 MiB, allow about 350 MiB for each state that parallel runs may save at once.
The tools show most refusals only as a status code (HTTP error: 409). The API logs every refusal of a state read, write, lock or unlock at WARNING, with its reason, the state, the lock IDs, and the run step or user that asked.
Network
The step container calls the API at HEGEMONY_SANDBOX_API_BASE_URL (default: HEGEMONY_API_BASE_URL). The worker resolves that host name itself and gives the address to the container, because containers on the sandbox cannot always resolve the platform's service names. It also lets the container reach that one endpoint through the flow's egress policy: an allow-listed policy gains it, and a deny-all policy allows only it. A deny rule that matches it still wins. A tf.plan whose State is backend or ephemeral gets no token and no route to the API. tf.apply takes its State from its plan step and always gets the route.