This page describes what each of Kumoss’s data stores holds: core-db,
phoenix-db, Redis, object storage, and the workspaces volume on the
filesystem shared between the core and the iac sidecar.
core-db
core-db (PostgreSQL 17) is the system of record for the application
domain. Thirteen tables: users (unique
on (issuer, subject), with email, display_name, operation_role,
nullable panel_role), sessions (owned by a user_id, including the
in_flight concurrency flag and the is_blocked compliance lock that
gates apply), workspaces, terraform_providers, pull_requests,
histories, statuses, rounds, and the artifact-metadata tables
artifacts, terraform_plans, reports, compliance_checks (the
audit verdict plus its stored findings), and code_changes. Roles are
managed from the admin panel or, for the first administrator on IdPs
that emit no email_verified claim, by a one-off SQL update (see
Grant the
first administrator access). There is
no migration tooling: the schema is created with SQLAlchemy
create_all, so a core_db_data volume from a pre-auth stack (sessions
keyed by a username column, no users table) must be migrated by
hand or recreated.
For the table-by-table breakdown, the enumerated columns, and how deletion cascades, see Database schema.
phoenix-db
phoenix-db (PostgreSQL 17) belongs exclusively to Phoenix and stores traces and prompt versions — operational/LLM telemetry, fully separate from Kumoss’s domain data.
Redis
Redis 8 is a cache, not a store: read-through/write-through for session facts, last statuses, and finished-session aggregates, with 2-second timeouts so a slow Redis fails open to PostgreSQL. It holds no locks and no pub/sub channels.
Object storage
Artifacts
Object storage holds generated artifacts under
sessions/<session id>/rounds/<round primary key>/, as
reports/<report type>-<token>.json, compliance/check-<token>.json,
plans/{plan|drift}-<token>.txt, or changes/<token>-<sanitized file name>.
RustFS is the default and bundled implementation: the Compose stack
starts it as the object-storage container, and because it is open
source (Apache-2.0) the default stack needs no proprietary or
source-available storage component. It is not the only option.
storage.provider selects one of three adapters, never more than one
at a time:
storage.provider |
Backend | Notes |
|---|---|---|
|
The bundled RustFS, or any S3-compatible object store reached through
a custom |
Presigned URLs are signed against |
|
AWS S3 on its regional endpoint |
Credentials from static keys or the AWS SDK default chain (instance role, workload identity) |
|
Azure Blob Storage in a storage account |
Shared-key authentication, SAS-token presigning; the account name is
derived from |
There is no dedicated Google Cloud Storage adapter; a GCS bucket could
only be reached through its S3-compatible interoperability endpoint
with the RUSTFS provider, which the repository does not test. For
state, a first-class gcs backend remains available by letting the
repository declare it — see below. Presign
expiry is 48 hours by default, floored at 30 hours so links outlive
cached session aggregates.
Terraform and OpenTofu state
Object storage for Terraform/OpenTofu state is on by default as
shipped. A remote backend, however, is not optional regardless of that
setting: workspaces are ephemeral and the .gitignore the core seeds
excludes .tfstate, so local state would be lost after the run.
Because storage.terraform_state_bucket ships set to
kumoss-terraform-state, the default is that *Kumoss manages state
itself, in a second bucket in the same store kept separate from the
artifacts bucket. A deployment can instead set the value to "" so
that each repository declares its own remote backend — azurerm,
s3, gcs, or anything else the engine supports — authenticated with
the cloud credentials given to the IaC sidecar, or hand the sidecar a
backend configuration file through IAC_BACKEND_CONFIG instead.
As shipped, the core creates that second bucket at boot and renders its
backend into a backend_override.tf in the workspace before every
init — so a repository’s own backend block is overridden without
editing its committed HCL, and the iac sidecar is the container that
must reach the store. Setting the value to "" turns all of that off:
no bucket is created, no override is written, and the workspace falls
back to whatever backend the repository’s own Terraform files declare.
A bare terraform_state_bucket: key with no value is not the same as
"" — it parses as YAML null and fails boot with a Pydantic validation
error — and deleting the key entirely also leaves managed state on,
since the configuration default is the same
kumoss-terraform-state value.
See Configure state backends for the full decision between the three ownership models and how core and sidecar configuration interact.
Workspaces volume
The workspaces volume is per-run: each run clones the repository into
a unique directory shared with the iac container, and the runner
removes it in a finally block. A successful generate round is the
exception — it renames its directory to /workspaces/<session>/pinned
before the cleanup runs, so that
directory survives on the volume until apply’s discard_pinned or
the next generate round replaces it. That pinned directory keeps the
initialized backend and the plan file, which is why apply executes the
reviewed plan without re-running init. Durable outputs otherwise leave
via git pushes and artifact uploads, not the volume.
authz sidecar persistence
authz persistence is a JSON file (/data/roles.json, container-local
in the shipped Compose file; set KUMOSS_AUTHZ_ROLE_STORE to a mounted
path to keep it), rewritten atomically under a process lock. It backs
the sidecar’s own user/role endpoints, which the core does not call;
Kumoss’s roles live in core-db.