This guide is for teams deploying Kumoss for shared organizational use. It explains what a production deployment consists of, which decisions apply to each component, and which settings and secrets you must supply. It does not repeat the field-by-field references; it points at them:
-
Configuration reference, for every
config.yamlfield. -
Environment variables and secrets, for every core and sidecar variable.
-
LLM providers and models, for LLM providers and credentials.
-
Enable authentication for authentication.
-
Monitor with Phoenix for tracing.
-
Customize prompts for the prompt registry.
-
HTTP APIs for the sidecar contracts in
contracts/openapi/and their conformance suites incontracts/conformance/.
Everything in this guide describes the implementation as it exists in the repository today. Where a capability is a recommendation rather than something the repository provides, it is labelled as such.
This entry page covers what production means, the reference topology,
the sidecars as integration boundaries, and the config.yaml
checklist. The credentials, networking, hardening, and operational
detail live on four child pages, each covering one step of the
deployment sequence.
At a glance
If you read nothing else, these are the production decisions this guide expands on:
| # | Decision | Where |
|---|---|---|
1 |
Run one core replica, one IaC sidecar per shared workspace volume; there is no multi-replica mode |
|
2 |
Enable OIDC; a blank issuer makes every caller a fully privileged user |
|
3 |
Treat the four sidecars as integration boundaries: implement IaC to your requirements, replace mapping and authorization or leave them disabled, integrate notifications |
|
4 |
Move every secret to a secret store; keep the variable names |
|
5 |
Use managed PostgreSQL, Redis, object storage (S3, Azure, or S3-compatible), and a persistent, access-controlled Phoenix; Phoenix is required at boot |
|
6 |
Reproduce the ingress rules: SSE path unbuffered, |
|
7 |
Give the IaC engine a scoped workload identity and container execution controls; decide who owns the Terraform state backend |
IaC (mandatory), below; Cloud credentials for the IaC engine; Configure state backends |
8 |
Review the seeded prompts before the first user session; they encode a generic policy |
|
9 |
Plan upgrades by hand: no database migrations, seeds never overwrite
prompts, |
The repository ships no Kubernetes manifests or Helm charts; the Compose file is the reference topology you translate to your platform.
What production means for Kumoss
Kumoss plans and applies infrastructure with real cloud credentials, pushes to your repositories, and sends every prompt, plan, and generated file to a tracing backend. A production deployment therefore has to control who can reach it, which cloud identity it acts with, where its data lives, and how it is operated.
Production is not the local Compose stack with environment:
production. That field only selects the Phoenix project prefix and the
prompt tag. It enables no authentication, no authorization, and no
stricter default.
The differences from Quickstart, in one table:
| Area | Local/non-production | Production |
|---|---|---|
Platform |
Docker Compose on one host |
A production platform such as Kubernetes, expressed by you from the reference topology below |
Authentication |
May stay disabled |
OIDC enabled; nothing else is acceptable because a blank issuer
makes every caller a |
Sidecars |
Bundled IaC, optionally bundled notifications |
All four treated as integration boundaries: reviewed, configured, hardened, replaced where needed |
Secrets |
Gitignored |
Kubernetes Secrets or an equivalent secret manager, injected as environment variables |
Data services |
Bundled PostgreSQL, Redis, RustFS, Phoenix on local volumes, with the well-known default credentials the samples ship |
Managed or operated services with persistence, backups, access control, and credentials from your secret store |
Terraform state |
Kumoss-managed state in the bundled RustFS by default, or each
repository’s own backend if you set |
A deliberate choice of owner: the repositories' own backends, a
Kumoss-managed bucket on |
Networking |
Plain HTTP on |
TLS, a real origin, an ingress that handles server-sent events, a protected Phoenix, a public object-storage endpoint |
The repository does not provide Kubernetes manifests, Helm charts,
Kustomize overlays, Terraform modules for the platform itself, database
migrations, or a multi-replica execution model. It provides the container
images, docker-compose.yml as the reference topology, and the OpenAPI
contracts and conformance suites for the sidecars; you write the platform
definition for your environment.
Reference topology
The Compose stack defines eleven services on one network. In production each row below is a deployment decision.
| Component | Reference image | Must be reachable by | Persistence | Production decision |
|---|---|---|---|---|
Core API |
|
Ingress ( |
Shared workspace volume |
Run one replica (see
Runtime
and scaling constraints). Inject |
Web application and edge |
|
Users |
none |
Keep it, or serve the built bundle from your own ingress. Reproduce its routing rules (Networking and TLS). |
PostgreSQL (core) |
|
Core |
Required |
Managed PostgreSQL recommended. Set |
Redis |
|
Core |
Not required (cache) |
Managed Redis or in-cluster. Must be reachable when the core boots. |
Object storage |
|
Core (SDK) and browsers (presigned URLs) |
Required |
Prefer AWS S3 or an Azure Storage Account through |
Phoenix |
|
Core (traces and prompt API), operators (UI) |
Required (through its PostgreSQL) |
Operate it with persistence and access control. It is required at core start-up. |
PostgreSQL (Phoenix) |
|
Phoenix |
Required |
Managed PostgreSQL recommended; set |
IaC sidecar |
|
Core only |
Shared workspace volume |
Mandatory. As a starting point — harden the image and container, or implement the contract yourself to your organization’s requirements. |
Mapping sidecar |
|
Core only |
none |
Replace with an implementation against your catalogue, or leave disabled. |
Notifications sidecar |
|
Core only |
none |
Use the bundled Slack implementation with a production webhook, or replace it. |
Authorization sidecar |
|
Core only |
none (container-local JSON role file; mount a path via
|
Replace with your policy implementation; the bundled one is permissive. |
All five Kumoss-built images (core, iac, mapping, notifications, authz)
run as the unprivileged kumoss user, uid/gid 10001. The
nginx/Dockerfile proxy image is the exception: it has no USER
directive and starts as root, dropping privileges only for its worker
processes; give it a non-root base image or a Kubernetes
runAsNonRoot/securityContext treatment if your policy requires the
master process itself to be non-root.
Two volumes matter beyond databases:
-
The shared workspace (
/workspacesin both the core and IaC containers,paths.upload_folder). The core clones repositories there and the IaC sidecar runs the engine against the same paths. It must be one filesystem mounted read-write by both containers at the same path, owned by uid/gid10001. On Kubernetes that means aReadWriteManyvolume or the two containers in the same pod sharing a volume, with a matchingfsGroup/runAsUser. Contents are ephemeral (each run’s directory is deleted afterwards), but the pinned plan of every session lives there between a generate round and its apply, so the volume must survive pod restarts. -
Phoenix’s database, which holds your curated prompts as well as traces. Losing it means re-seeding from the repository defaults and losing every prompt edit.
Sidecars as production integration boundaries
Each sidecar is an OpenAPI contract. The core calls it with a bearer
token from the environment variable named by services.<name>.token_env,
and the sidecar validates that token. Any implementation of the
contract can be used by changing services.<name>.endpoint. Use the
conformance suites in contracts/conformance/ to verify a
replacement; every page under HTTP
APIs carries the command for its own sidecar.
Common rules for every enabled sidecar:
-
Set a distinct, random bearer token (for example
openssl rand -hex 32) on both sides. The core refuses to boot when the IaC sidecar’s or an enabled optional sidecar’s token variable is empty. On the sidecar side, an empty token disables the check entirely, which is never acceptable in production. -
The bundled sidecars compare tokens in constant time, but they remain simple reference services. Keep every sidecar on a private network segment reachable only by the core, and never expose one through the ingress.
-
The
services.<name>.timeoutfield is honoured by the IaC and notifications clients; the mapping and authorization clients use fixed budgets of 10 and 15 seconds. -
Every sidecar exposes
GET /healthzfor liveness probes.
IaC (mandatory)
The bundled executor is a starting point, not production-ready: harden the image and container (the checked-in compose applies no hardening), review the credential pass-through, or implement the contract yourself. It is functionally complete (it runs every command the core needs), but it encodes no organizational policy: it executes provider plugins from generated HCL with whatever identity is in its environment, on a shared filesystem, with no approval step, no engine timeout, and in-memory job state.
For production the recommendation is to implement the IaC contract
(contracts/openapi/iac.v1.yaml)
according to your organization’s
requirements, not only to harden the bundled image. Typical reasons:
running the engine on your existing Terraform or OpenTofu execution
platform, using a workspace and state strategy other than a shared
volume, enforcing per-project cloud identities, adding policy or
approval hooks around plan and apply, auditing every command, or
supporting clouds and CLIs the bundled image does not include. The
contract is small (one command per asynchronous job, polled by the
core), and the
conformance suite in
contracts/conformance/iac/ verifies a replacement.
If you nevertheless start from the bundled image, its identity and execution environment are the security decisions you must make:
-
Cloud identity. The engine reads whatever credentials the provider supports from the container environment (
ARM_*,GOOGLE_*,AWS_*, mounted files). Use a workload identity (Kubernetes service account federation, pod identity, instance roles) scoped to the projects Kumoss may change, rather than static keys. Missing credentials do not stop the container; the engine fails atplanorapplywith its own error. -
Execution controls. The image already runs as an unprivileged user. The checked-in Compose file applies nothing else. For production, run the container with all Linux capabilities dropped,
no-new-privileges, a read-only root filesystem with writable/tmpand/home/kumoss, and CPU, memory, and process limits. These were validated on a feature branch of this repository but are not part of the checked-in Compose file; treat them as a requirement you implement on your platform. -
No engine timeout. The sidecar does not bound the engine process itself;
services.iac.job_timeoutin the core (default one hour per job) is the only bound. Set it deliberately. -
Egress.
initdownloads providers fromregistry.opentofu.org(orregistry.terraform.iowhenIAC_BINARY=terraform), andplan/applyreach your cloud APIs. Corporate TLS inspection of the registry host breaks provider downloads unless the CA is mounted into the container. -
Jobs are in memory. A restart forgets queued and finished jobs; run one instance per shared workspace and do not scale it horizontally.
-
Engine choice.
IAC_BINARY=tofu(default) runs OpenTofu;terraformruns the bundled HashiCorp Terraform 1.16.0, which is BUSL-1.1 licensed and makes your use subject to its terms. -
State backend. The sidecar is always the process that executes the backend, so it always needs reach to it and credentials for it — whichever model you pick. By default (
storage.terraform_state_bucketset tokumoss-terraform-state, as shipped) the core configures it: state goes to that bucket in the same object store as artifacts, keyed per project, and the sidecar needs network reach tostorage.endpoint_urlplus, where the rendered block omits static keys, its own credentials for the bucket. Setstorage.terraform_state_bucketto""to leave the backend to each target repository instead, with no Kumoss configuration involved. For one central backend defined outside Kumoss, mount a file and pointIAC_BACKEND_CONFIGat it. All three models: Configure state backends. -
Blast radius of the sidecar’s environment.
KUMOSS_IAC_TOKENand every other variable in the IaC sidecar’s environment are reachable by the provider plugins the generated HCL invokes, and by anylocal-execprovisioner in that HCL — generated code runs with the same privileges as the sidecar process. The job’sworkspace_pathis not restricted to/workspacesby the sidecar itself. Treat the sidecar as reachable only from the core and give it nothing more than the cloud identity it needs.
Mapping (optional)
The bundled implementation is a direct passthrough for https://
identifiers: the value the user types is returned as the repository URL.
Unlike leaving the integration disabled, which echoes every identifier,
the bundled service answers 404 for anything else. Enable it only with
an implementation that resolves your business identifiers (CMDB, Backstage, a catalogue)
against the canonical IaC repository. Once enabled, a slow or
unreachable sidecar makes the wizard’s resolve step fail with
504/502.
Your implementation may also answer terraform_provider and scope_id
(the cloud subscription, project, account, compartment, or cluster).
Both are optional and both are a promise: the wizard stops asking the
user for any field your service fills in, so a value it cannot stand
behind is a wrong answer nobody gets the chance to correct. Return
null — meaning "I do not know" — whenever the catalogue is silent, and
never infer a provider from the shape of the repository URL.
Notifications (optional, recommended)
The bundled implementation posts to one Slack incoming webhook and is
suitable for production once the webhook and token come from your
secret store. The core emits iac.compliance.failed,
iac.impact.high, iac.apply.failure, and system.exception.failure,
plus user-originated support requests from the web application; the
audience is the session owner (or caller) plus every user with panel
role editor or higher. If your operational channel is not Slack
(Teams, e-mail, PagerDuty, an ITSM tool), implement the contract in
contracts/openapi/notifications.v1.yaml.
Delivery is best-effort: a
failure is logged and never fails a session.
Authorization (optional, must be replaced if enabled)
The bundled implementation answers authorized: true to every
cloud-project check (KUMOSS_AUTHZ_PERMISSIVE=true) or authorized:
false to every check (any other value). Neither is a policy. Do not
enable the bundled authorization sidecar as your production
access-control policy.
Understand what this integration is and is not:
-
It is consulted only by the wizard’s preflight
POST /api/v1/auth/authorize, which sends the cloud, project name, environment, and the caller’s e-mail (or OIDC subject) toPOST /v1/check. It answers whether that user may work on that cloud project. -
It is not Kumoss’s own access control. Who can log in, which operations a user may run, and who can see or act on a session are decided by OIDC plus the operation and panel roles stored in the core database (Administer users and locks).
-
While disabled, the core answers "authorized" without calling anything. Once enabled, an unreachable sidecar produces a
502/504error in the wizard, not an allow.
Implement the contract in
contracts/openapi/authz.v1.yaml
against your entitlement source and enable it.
Configuration checklist (config.yaml)
config.yaml is baked into the core image at build time, or mounted at
the path in KUMOSS_CONFIG. It contains no secrets and can live in
version control. Review every section; the
configuration reference has the
fields and a
production-shaped
example.
| Section | Production requirement |
|---|---|
|
|
|
|
|
|
|
Endpoint of your IaC sidecar, token variable name, |
|
Enabled only with a production-grade implementation behind |
|
Keep |
|
Two model strings (or |
|
|
|
Variable names only; values are secrets. |
|
Your Phoenix base URL, ending in |
|
Your real origin, for example |
|
Provider, artifacts bucket or container, the endpoint the core calls,
and the public endpoint browsers reach. Also |
|
Provider matching your repository host, author identity. |
Changing any of these means rebuilding the core image (or updating the mounted file) and restarting the core.