Kumoss

Data and state

What core-db, phoenix-db, Redis, object storage, and the workspaces volume each hold.

This page describes what each of Kumoss’s data stores holds: core-db, phoenix-db, Redis, object storage, and the workspaces volume on the filesystem shared between the core and the iac sidecar.

core-db

core-db (PostgreSQL 17) is the system of record for the application domain. Thirteen tables: users (unique on (issuer, subject), with email, display_name, operation_role, nullable panel_role), sessions (owned by a user_id, including the in_flight concurrency flag and the is_blocked compliance lock that gates apply), workspaces, terraform_providers, pull_requests, histories, statuses, rounds, and the artifact-metadata tables artifacts, terraform_plans, reports, compliance_checks (the audit verdict plus its stored findings), and code_changes. Roles are managed from the admin panel or, for the first administrator on IdPs that emit no email_verified claim, by a one-off SQL update (see Grant the first administrator access). There is no migration tooling: the schema is created with SQLAlchemy create_all, so a core_db_data volume from a pre-auth stack (sessions keyed by a username column, no users table) must be migrated by hand or recreated.

For the table-by-table breakdown, the enumerated columns, and how deletion cascades, see Database schema.

phoenix-db

phoenix-db (PostgreSQL 17) belongs exclusively to Phoenix and stores traces and prompt versions — operational/LLM telemetry, fully separate from Kumoss’s domain data.

Redis

Redis 8 is a cache, not a store: read-through/write-through for session facts, last statuses, and finished-session aggregates, with 2-second timeouts so a slow Redis fails open to PostgreSQL. It holds no locks and no pub/sub channels.

Object storage

Artifacts

Object storage holds generated artifacts under sessions/<session id>/rounds/<round primary key>/, as reports/<report type>-<token>.json, compliance/check-<token>.json, plans/{plan|drift}-<token>.txt, or changes/<token>-<sanitized file name>. RustFS is the default and bundled implementation: the Compose stack starts it as the object-storage container, and because it is open source (Apache-2.0) the default stack needs no proprietary or source-available storage component. It is not the only option. storage.provider selects one of three adapters, never more than one at a time:

storage.provider Backend Notes

RUSTFS (default)

The bundled RustFS, or any S3-compatible object store reached through a custom endpoint_url (path-style addressing, static access and secret keys)

Presigned URLs are signed against public_endpoint_url, the address browsers reach

S3

AWS S3 on its regional endpoint

Credentials from static keys or the AWS SDK default chain (instance role, workload identity)

STORAGE_ACCOUNT

Azure Blob Storage in a storage account

Shared-key authentication, SAS-token presigning; the account name is derived from endpoint_url and the container from bucket

There is no dedicated Google Cloud Storage adapter; a GCS bucket could only be reached through its S3-compatible interoperability endpoint with the RUSTFS provider, which the repository does not test. For state, a first-class gcs backend remains available by letting the repository declare it — see below. Presign expiry is 48 hours by default, floored at 30 hours so links outlive cached session aggregates.

Terraform and OpenTofu state

Object storage for Terraform/OpenTofu state is on by default as shipped. A remote backend, however, is not optional regardless of that setting: workspaces are ephemeral and the .gitignore the core seeds excludes .tfstate, so local state would be lost after the run. Because storage.terraform_state_bucket ships set to kumoss-terraform-state, the default is that *Kumoss manages state itself, in a second bucket in the same store kept separate from the artifacts bucket. A deployment can instead set the value to "" so that each repository declares its own remote backend — azurerm, s3, gcs, or anything else the engine supports — authenticated with the cloud credentials given to the IaC sidecar, or hand the sidecar a backend configuration file through IAC_BACKEND_CONFIG instead.

As shipped, the core creates that second bucket at boot and renders its backend into a backend_override.tf in the workspace before every init — so a repository’s own backend block is overridden without editing its committed HCL, and the iac sidecar is the container that must reach the store. Setting the value to "" turns all of that off: no bucket is created, no override is written, and the workspace falls back to whatever backend the repository’s own Terraform files declare. A bare terraform_state_bucket: key with no value is not the same as "" — it parses as YAML null and fails boot with a Pydantic validation error — and deleting the key entirely also leaves managed state on, since the configuration default is the same kumoss-terraform-state value.

See Configure state backends for the full decision between the three ownership models and how core and sidecar configuration interact.

Workspaces volume

The workspaces volume is per-run: each run clones the repository into a unique directory shared with the iac container, and the runner removes it in a finally block. A successful generate round is the exception — it renames its directory to /workspaces/<session>/pinned before the cleanup runs, so that directory survives on the volume until apply’s discard_pinned or the next generate round replaces it. That pinned directory keeps the initialized backend and the plan file, which is why apply executes the reviewed plan without re-running init. Durable outputs otherwise leave via git pushes and artifact uploads, not the volume.

authz sidecar persistence

authz persistence is a JSON file (/data/roles.json, container-local in the shipped Compose file; set KUMOSS_AUTHZ_ROLE_STORE to a mounted path to keep it), rewritten atomically under a process lock. It backs the sidecar’s own user/role endpoints, which the core does not call; Kumoss’s roles live in core-db.