prerelease Prerelease stable Latest
Kumoss
prerelease Prerelease stable Latest

Configure state backends

Choose which party owns Terraform and OpenTofu state — the repository, Kumoss, or the deployment — and understand which value wins when more than one is set.

Kumoss’s workspaces are ephemeral by design: every call clones the repository into a fresh directory on the shared workspaces volume and deletes it afterwards. A local terraform.tfstate dies with that directory, and the next session plans against empty state, proposing to create infrastructure that already exists. Every project therefore needs a remote backend, and this page helps you decide who provides it.

The seeded .gitignore excludes .tfstate, .tfstate.*, and *_override.tf, so state and the generated backend file never reach a commit or a pull request.

Choose a model

Exactly one party decides where state goes. Answer these questions in order and stop at the first yes:

  1. Does every repository already declare a backend that must keep working (shared with CI or with engineers' laptops)? Use Repository-declared backend.

  2. Must one central backend, defined by the platform team, serve every repository? Use Sidecar-supplied backend.

  3. Otherwise, keep the shipped default, Kumoss-managed state: the repositories need no backend block at all.

Two settings select the model, one on the core and one on the IaC sidecar:

storage.terraform_state_bucket (core config.yaml) IAC_BACKEND_CONFIG (services/iac/.env) Effective model Who decides where state goes

a bucket name (shipped)

unset (shipped)

Kumoss-managed

The core, from storage.*

""

unset

Repository-declared

Each repository’s own backend block

""

a file path

Sidecar-supplied

The deployment, through a mounted file

a bucket name

a file path

Mixed — a misconfiguration

Nobody fully; see Which value wins

The shipped default: the core creates the bucket and writes the backend for you.

Blank the state bucket and let each repository’s own backend block own state.

Mount a backend configuration file on the IaC sidecar and own the decision centrally.

Turn managed state off

Kumoss-managed state is on as shipped, so the two other models start by turning it off. Only an explicit empty string does that:

storage:
  terraform_state_bucket: ""    # "" or whitespace-only: managed state off

The other spellings do not behave the same way:

What you write Result

terraform_state_bucket: ""

Managed state off.

The key deleted

The field default kumoss-terraform-state applies: managed state stays on.

terraform_state_bucket: with nothing after it

YAML null: the core fails at boot with storage.terraform_state_bucket / Input should be a valid string.

config.yaml is baked into the core image, so rebuild after the change (docker compose build core && docker compose up -d core), or mount the file and point KUMOSS_CONFIG at it, as described in Images and configuration delivery.

Which value wins

When several layers configure the backend, the engine merges them key by key. Highest priority first:

  1. IAC_BACKEND_CONFIG: the file the sidecar passes to init as -backend-config=<file>.

  2. backend_override.tf: written by the core when managed state is on.

  3. The repository’s own terraform { backend } block.

A key that a higher layer does not set falls through to the next one. The backend type (azurerm, s3, …​) always comes from the workspace, either from the repository or from the override.

Example. A repository declares a backend, managed state is left on, and the sidecar also mounts a file:

# 3. The repository's main.tf
terraform {
  backend "azurerm" {
    storage_account_name = "acmetfstate"
    container_name       = "tfstates"
    key                  = "platform.tfstate"
  }
}

# 2. backend_override.tf, written by the core (managed state on)
terraform {
  backend "azurerm" {
    storage_account_name = "acmekumoss"
    container_name       = "kumoss-terraform-state"
    key                  = "<project_id>/terraform.tfstate"
    access_key           = "..."
  }
}

# 1. The file IAC_BACKEND_CONFIG points at
key = "central.tfstate"

What init binds:

Key Effective value Taken from

key

central.tfstate

The sidecar file

storage_account_name

acmekumoss

The core’s override

container_name

kumoss-terraform-state

The core’s override

access_key

the core’s key

The core’s override

Two lessons follow:

  • A repository backend is ignored whenever managed state is on. None of the repository’s values survive, and Kumoss logs no warning about it. The plan runs against whatever the managed key holds, usually nothing. Turn managed state off to keep a repository’s backend.

  • Never combine managed state with IAC_BACKEND_CONFIG. The result is a backend that neither side fully describes. A file that supplies a single key also removes Kumoss’s per-project key, so every project shares one state object.

The order was verified empirically by running tofu init and terraform init against a local backend with the same key set in all three layers.

Change models safely

Switching models moves no state. Turning managed state on does not import the state your repositories' backends hold; turning it off leaves what Kumoss stored in its bucket where those backends do not look.

init always runs with -reconfigure (tofu init -no-color -input=false -reconfigure), because -input=false cannot answer the "Backend configuration changed" prompt. As a result:

  • State already in the new backend is adopted. Re-pointing a project at a backend that already holds its state is safe.

  • State under the previous backend is not migrated. Kumoss never copies state between backends.

  • A project starts from empty state after any change to its backend or its key: switching storage.provider, renaming terraform_state_bucket, moving the root module, or changing the cloud scope. The next plan proposes creating resources that already exist.

Migrate before the next session runs, with the engine itself:

tofu state pull > project.tfstate           # against the old backend
tofu init -reconfigure -backend-config=...  # point at the new backend
tofu state push project.tfstate

tofu init -migrate-state, run by hand in an equivalent workspace, works too. Keep the pulled file as a backup, and verify with a plan that shows no changes before letting Kumoss run again.

What Kumoss does not do

State ownership stops at "write the backend and key". Everything below is yours:

  • No state migration, between any two backends (see above).

  • No versioning, backups, or retention. Kumoss does not enable versioning, object lock, soft delete, or point-in-time restore, and it never deletes a state object.

  • No DynamoDB lock table. Locking uses the S3-native lockfile or the Azure blob lease; there is nothing to provision.

  • No gcs backend adapter. storage.provider has three members (RUSTFS, S3, STORAGE_ACCOUNT). For a gcs backend, use the repository-declared or sidecar-supplied model and let the repository’s block, or a mounted file, declare gcs.

  • No state encryption beyond the store’s own. State can contain secrets; rely on encryption at rest and restrict access accordingly.

  • Import writes to state during the round. The import modes run terraform import before any pull request is reviewed. See Import infrastructure.

Troubleshooting

Find the active state

Start here whenever you are unsure which state file a project uses. The answer depends on three layers (see Which value wins), so check all three, highest priority first. The commands read the running containers, not your local files, so they show what is actually deployed.

  1. Sidecar layer. Check whether the IaC sidecar overrides the backend:

    docker compose exec iac sh -c 'echo "IAC_BACKEND_CONFIG=${IAC_BACKEND_CONFIG:-<unset>}"; [ -n "$IAC_BACKEND_CONFIG" ] && cat "$IAC_BACKEND_CONFIG"'
    IAC_BACKEND_CONFIG=<unset>

    A path that does not start with / is relative to each cloned workspace, so cat fails here even when init finds the file. Read that file in the repository instead.

  2. Core layer. Check whether Kumoss-managed state is on in the image that is running:

    docker compose exec core python -c "
    from src.shared.config import system_config as c
    s = c.storage
    print('managed state :', 'ON' if s.state_bucket else 'OFF')
    print('provider      :', s.provider.name)
    print('bucket        :', s.state_bucket or '-')
    print('endpoint      :', s.endpoint_url)
    "
    managed state : ON
    provider      : RUSTFS
    bucket        : kumoss-terraform-state
    endpoint      : http://object-storage:9000
  3. Repository layer. When managed state is OFF, read the backend block the repository commits:

    git -C _REPOSITORY_PATH_ grep -n -A8 'backend "'
    main.tf:2:  backend "azurerm" {
    main.tf-3-    resource_group_name  = "tfstate-rg"
    main.tf-4-    storage_account_name = "acmetfstate"
    main.tf-5-    container_name       = "tfstates"
    main.tf-6-    key                  = "platform.tfstate"

Read the results together:

Sidecar layer Core layer The active state is…

<unset>

ON

<bucket>/<project_id>/terraform.tfstate in the store at endpoint. Get the project_id from the core log (see below).

<unset>

OFF

Wherever the repository’s backend block points. In the example above: container tfstates, blob platform.tfstate in account acmetfstate.

a path

OFF

The repository’s backend block, with every key the file sets replacing the repository’s value.

a path

ON

Mixed: a misconfiguration. Fix it before trusting any state (see Which value wins).

Get the project_id under managed state. The core logs the bucket and key each time it writes the override, before every init:

docker compose logs core | grep "Terraform state:"
Terraform state: rustfs bucket=kumoss-terraform-state key=3f9a…c41e/terraform.tfstate

Use this line to find the key, not to decide which model is active. It is written by the core, so:

  • It appears only when managed state is on. No line does not tell you where state is, only that the repository or the sidecar decides.

  • It does not include IAC_BACKEND_CONFIG overrides, which the core never sees.

  • It prints no account or endpoint; those come from storage.endpoint_url.

  • It appears only after a session has run init since the container was last created, and an apply logs nothing new: it reuses the workspace the plan initialized, so it writes to the same state.

Confirm the object exists and is the one being written. After an apply, the state object’s modification time must change:

az storage blob show --auth-mode login \
  --account-name acmetfstate --container-name tfstates \
  --name platform.tfstate \
  --query "{size: properties.contentLength, modified: properties.lastModified}"
aws s3api head-object --bucket acme-team-tfstate \
  --key envs/prod/terraform.tfstate \
  --query "{size: ContentLength, modified: LastModified}"
# One-off: the backend uses path-style addressing, so the client must too.
aws configure set default.s3.addressing_style path

AWS_ACCESS_KEY_ID=rustfsadmin AWS_SECRET_ACCESS_KEY=rustfsadmin \
AWS_DEFAULT_REGION=us-east-1 \
  aws --endpoint-url http://localhost:9000 \
  s3 ls s3://kumoss-terraform-state/ --recursive

Plans against empty state

Symptom Cause and fix

The repository declares a backend, yet every plan proposes creating resources that already exist

Managed state is on and its override replaces the repository’s backend (see Which value wins). The log line above appears. Set storage.terraform_state_bucket: "" and rebuild the core image.

Every session plans against empty state, and nothing is ever recorded

Managed state is off, but the repository declares no backend: init succeeds against the local backend, and the state dies with the clone. Add a terraform { backend …​ } block, or turn managed state back on.

Plans propose recreating everything after a configuration change

The project’s state key or backend changed and the state was not migrated. Compare the logged key= with the bucket contents, then migrate (see Change models safely).

Two repositories collide on one state file

An IAC_BACKEND_CONFIG file supplies a single key for every repository. Declare distinct keys per repository, or use managed state alone.

Boot and init failures

Symptom Cause and fix

Core fails at boot: Failed to initialize object storage

Managed state is on and the state bucket cannot be created or reached (credentials, CreateBucket permission, endpoint). The artifact and state buckets are checked together; verify both, or pre-create them.

init fails with a credentials or AccessDenied error

The IaC sidecar lacks credentials for the state store, or its identity has no access to it. The core’s credentials do not count. See the two credential grants and Credentials required.

init on azurerm cannot find the storage account

init runs with ARM_SUBSCRIPTION_ID set from the request’s cloud scope. Pin subscription_id in the backend block when the state account lives in another subscription.

init fails with InvalidAccessKeyId after switching to provider: S3

RUSTFS_ACCESS_KEY / RUSTFS_SECRET_KEY are still rustfsadmin in core/.env and were embedded in the override. Blank both and rebuild.

init fails to dial the endpoint (connection refused, DNS failure)

The sidecar cannot reach the state store. Put it on the same network (bridge-network in Compose) or allow that egress.

init fails on the backend configuration file

IAC_BACKEND_CONFIG is not checked at startup: the path is wrong, the file is not mounted, or uid 10001 cannot read it. The error is in the failing job’s stderr.

Error acquiring the state lock

Another run holds the lock, or a crashed run left it stale. If stale, run tofu force-unlock LOCK_ID in an equivalent workspace, or remove the .tflock object or break the blob lease.

Session fails with a Terraform backend error before init runs

The core could not write backend_override.tf. The workspaces volume must be writable by uid/gid 10001 (chown -R 10001:10001 on a volume from a root-running stack).

The override file in a pull request

If a backend_override.tf shows up in a pull request, the workspace .gitignore write failed: the core logs terraform gitignore couldn’t be created, carries on, and git add -A picks the file up. Check that the workspaces volume is writable by uid/gid 10001, remove the file from the branch, and re-run the round.

A repository without a .gitignore is not the cause: the seeded file is honored for the working tree even though it is never committed. A repository that already tracks a .gitignore gets Kumoss’s template appended, and that change is committed to the session branch.