This is the last of four pages that expand Deploy to production: with networking and hardening decided, this page covers what each data service holds and how to run it in production, how to back up and upgrade a deployment, how to verify a fresh deployment, and how to diagnose the failures a production deployment is most likely to hit.
Data services and persistence
PostgreSQL (core). The system of record: users and roles, sessions
and their ownership, rounds, statuses, conversation history, pull
requests, and artifact metadata. Use a managed instance with backups
and TLS. There are no migrations: the schema is created with
create_all at start-up, and a schema change between versions requires
a manual migration or a fresh database. Plan upgrades accordingly.
Redis. A cache; the database remains authoritative and cache operations fail open at runtime. The boot-time connectivity check must succeed, so Redis has to exist, but it needs no persistence.
Object storage. storage.bucket holds reports, plans, drift output,
and code-change artifacts under
sessions/<session>/rounds/<round>/. A second bucket,
storage.terraform_state_bucket, holds Terraform/OpenTofu state under
<project_id>/terraform.tfstate — it ships set to
kumoss-terraform-state, so this is on by default; set it explicitly
to "" to leave state to each repository’s own backend instead. Every
configured bucket is created at boot if missing, and the core fails to
boot if the store is unusable. Choose storage.provider:
-
S3: AWS S3 on its regional endpoint, credentials from the default chain or static keys. State uses thes3backend withuse_lockfile = true. -
STORAGE_ACCOUNT: an Azure storage account with a shared key. State uses theazurermbackend with native blob leases. -
RUSTFS: RustFS or any S3-compatible server with a custom endpoint.
Browsers download artifacts through presigned URLs built against
storage.public_endpoint_url, so that endpoint must be reachable from
users' browsers over TLS, and the bucket policy must allow presigned
reads. Links are valid for storage.presign_expiry_seconds (48 hours
by default). Apply lifecycle and retention rules; Kumoss never deletes
stored artifacts (it only removes an object whose database record
failed to be written).
Since Kumoss-managed state is on by default, treat the state bucket differently from the artifacts bucket. It is never browser-facing and needs no presigned reads; it must be reachable from the IaC sidecar, not from users; and it should carry versioning and soft delete, be excluded from any artifact expiry rule, and be included in your restore drills — it is the record of what Kumoss has built. Kumoss never deletes a state object. Details and per-provider IAM minimums: Operations.
Phoenix. Required at core start-up (prompt seeding retries for about
27 seconds and then aborts the boot) and on every request (prompt
fetch). Run it with its own PostgreSQL, back that database up, and put
the UI and API behind access control: Phoenix has no authentication of
its own, and it holds prompts, plans, generated code, repository
metadata, and user identifiers. The core sets no exporter headers; OTLP
header env vars may be honoured by the OpenTelemetry SDK but this is
not verified here — put an authenticating proxy or collector in front
of Phoenix if you need auth. Note that the same
telemetry.collector_url serves both trace export and the prompt API
(Configuration).
Checked-in defaults you must not carry into production. The compose file hard-codes several credentials and settings that are fine only on an isolated workstation:
-
POSTGRES_PASSWORD=postgreson bothcore-dbandphoenix-db, and the matching password embedded inPHOENIX_SQL_DATABASE_URL. Replace all three with a generated password from your secret store, and updateKUMOSS_SQL_DATABASE_URLandPHOENIX_SQL_DATABASE_URLto match. -
object-storagesets noRUSTFS_ACCESS_KEY/RUSTFS_SECRET_KEY, so RustFS runs with its built-inrustfsadmin/rustfsadmincredentials. Before onboarding real data, set both onobject-storageto generated values and set the same values in the core’sRUSTFS_ACCESS_KEY/RUSTFS_SECRET_KEY. Set them explicitly on both sides: when the variables are absent, RustFS and the core (withstorage.providerset toRUSTFS) both silently fall back torustfsadmin. Deleting them does not remove the default; only setting them does. -
Floating image tags (
rustfs/rustfs:latest,ghcr.io/astral-sh/uv:latestin every Python Dockerfile,node:24-alpine,postgres:17,redis:8.8) are convenient for local use but not reproducible. Pin every image you run in production by digest.
RustFS console. Kumoss does not ship or configure the RustFS web
console. RustFS turns it on by default, so the compose file sets
RUSTFS_CONSOLE_ENABLE=false on object-storage. Kumoss only needs the
S3 API on port 9000, and the nginx reference proxies nothing else. If
you want the console to browse buckets or manage access keys, set it up
yourself by following the
RustFS console
documentation. With the bundled compose file that means:
-
Set
RUSTFS_CONSOLE_ENABLE=trueonobject-storage. The console listens onRUSTFS_CONSOLE_ADDRESS,:9001by default, and is served under/rustfs/console/. -
Make port 9001 reachable. Nothing publishes or proxies it today. For a workstation, bind it to loopback only:
object-storage: environment: - RUSTFS_CONSOLE_ENABLE=true ports: - "127.0.0.1:9001:9001"Then open
http://localhost:9001/rustfs/console/and sign in with the RustFS credentials:rustfsadmin/rustfsadminunless you setRUSTFS_ACCESS_KEY/RUSTFS_SECRET_KEYonobject-storage. -
Beyond a workstation, follow the RustFS recommendations: serve the console over TLS, restrict network access to its listener, rotate the
rustfsadminkeys before anyone else can reach it, and use least-privilege accounts for routine work. Never put the console on the public presigned-URL endpoint (port 9000). The image also shipsRUSTFS_CONSOLE_CORS_ALLOWED_ORIGINS=*, so scope that to your console’s origin if you use it cross-origin.
Backups and upgrades
Back up: the core PostgreSQL database, the Phoenix PostgreSQL database (prompts and traces), the artifacts bucket, and — with the highest priority — wherever Terraform state lives, since losing it means losing the record of the infrastructure Kumoss manages. That is the state bucket if you enabled Kumoss-managed state, and otherwise the backend each repository declares, which is covered by whoever owns it. The workspace volume needs no backup beyond surviving restarts.
Upgrading Kumoss:
-
Read the release notes for schema changes; there are no automatic migrations. One known breaking change predates OIDC support and is documented in full in Data and state: migrate or drop the affected volume by hand before you point a current core at it.
-
Rebuild or pull the images and the configuration for the new version.
-
New prompt seed files are created in Phoenix on the next core start; existing prompts are never overwritten, so review the seed changes and apply the ones you want as new prompt versions in Phoenix.
-
Do not change
environmenton an existing deployment without first tagging every prompt version with the new value.
Verification checklist
After the first deployment:
# Public route: issuer must be non-empty in production
curl https://kumoss.example.invalid/api/v1/auth/config
# Protected route: must answer 401 without a token
curl -i https://kumoss.example.invalid/api/v1/users/me
# Sidecar liveness from inside the cluster network
curl http://iac:8082/healthz
Then, as the bootstrap administrator:
-
Sign in; confirm the user page shows the Admin section and the admin portal lists users.
-
Run a generate session against a test repository and a test cloud project; confirm the branch, the pull request, the report, and the traces in Phoenix under
pro-terraform-day2. Also confirm state landed where you intended: with the shipped default (Kumoss-managed state), the core logsTerraform state: <provider> bucket=<bucket> key=<project_id>/terraform.tfstatebeforeinitand that object must exist in the state bucket afterwards; if you setstorage.terraform_state_bucketto"", check the backend the repository declares instead, and confirminitdid not silently fall back to local state. Either way the pull request must contain nobackend_override.tfand no*.tfstate. -
Trigger a compliance failure or a high-impact change and confirm the session is locked and the notification arrives.
-
Apply from a session that passed, and confirm the plan that ran is the one reviewed.
-
Restart the core and confirm it boots (prompt seeding reports
already present) and that sessions resume normally.
Troubleshooting
| Symptom | Cause | Fix |
|---|---|---|
Core exits with |
The IaC sidecar’s token variable, or that of a sidecar enabled in
|
Set a distinct bearer token on both sides of every enabled sidecar. |
Core exits with |
Phoenix is not reachable at |
Fix the URL or the network path, then restart the core. |
Core boots with every optional sidecar disabled although
|
|
Fix the mount path so |
Users get |
Usually the API scope or |
See OIDC troubleshooting. |
The session view never updates |
The ingress buffers the SSE path. |
Disable buffering for |
Artifact links fail in the browser |
|
Point |
Every plan fails with |
The shared workspace is not writable by uid |
|
|
Egress to |
Allow the egress, or mount your inspection CA into the IaC container. |
Every session plans against empty state and nothing is recorded |
This happens when a deployment has explicitly set
|
Add a backend to the repository, or remove the explicit |
|
|
Blank both and rebuild the core image. |
|
The IaC sidecar lacks credentials or permissions for the state
bucket, or cannot reach |
See Credentials required. |
Every plan proposes creating resources that already exist |
The project’s state key changed, or state was not migrated after a
backend change; |
|
|
Another run holds it, or a crashed run left it stale. |
Clear it deliberately (Locking). |
|
|
Fix the mount and its ownership, or unset the variable. |
A run stopped without a failure message |
A prompt was missing in Phoenix for the deployment’s |
See Runtime lookup. |