Kumoss calls large language models (LLMs) through
LiteLLM, a routing library that speaks to
many providers behind one interface. This page explains how Kumoss
selects models, where credentials go, what the core checks at startup,
and how to use the advanced llm.model_list router configuration.
All values in the examples are fictitious. Domains use .invalid, and
keys look like sk-example-not-a-real-key. Never paste real secrets
into config.yaml or into documentation.
Related guides: Configuration
(every config.yaml field), Environment
variables and secrets (the .env files), and the two installation
guides, Quickstart and
Deploy to production, which
say where the credentials live in each deployment model.
How Kumoss uses LiteLLM
Kumoss configures two model roles in config.yaml. Both are applied by
the core’s LiteLLM Router; there is no provider enum in Kumoss, so any
model string LiteLLM understands is accepted by the configuration.
Whether a model works is a different matter; see Model
requirements below.
| Key | Role | Used for |
|---|---|---|
|
High-quality model |
The IaC generator chain (writing and fixing Terraform code), the target generator chain (choosing plan targets), the report generator chain, and the task splitter chain (turning a drift report into remediation operations). |
|
Fast, cheaper model |
Every other chain: request filtering, drift reconciliation filtering, drift exception filtering, prompt composition, pull-request text, the compliance checker, status messages, and the LLM calls that some tools make internally (for example web search). |
Three extra settings apply to both roles:
| Key | Default | Meaning |
|---|---|---|
|
|
Sampling temperature sent with every chat completion. When a chain
requests extended thinking, the core sends |
|
|
Sent as |
|
|
Seconds LiteLLM waits for a single inference call before aborting it. Each of the three internal retries gets a fresh budget, so the worst case for one inference is four times this value. Raise it for slow reasoning models, lower it to fail faster on a stalled provider. |
The core passes drop_params=True to LiteLLM, so a provider that does
not support one of these parameters ignores it instead of failing.
Model requirements
Kumoss’s agents are tool-calling loops, so the configuration accepting a model string does not mean the model can run a session. Both roles must provide:
-
Reliable OpenAI-style function calling. Agent loops send a
toolslist withtool_choice: autoand expect well-formed JSON arguments back; each loop ends only when the model calls its sentinel tool. A reply without tool calls is answered with a reminder and costs a turn of the loop budget (see Agent tool loop), so a model that often answers in text instead of calling tools runs out of turns. This applies to agent loops specifically — the status/summary text calls (generate_text) send no tools at all. Models without reliable function calling (many small self-hosted models) fail every agent session. -
A completion cap of at least
llm.max_output_tokens(32000 by default), or lower that setting to what the provider allows. -
A finish reason of
tool_useorend_turn. Any other finish reason —length(the completion was truncated),content_filter, and so on — aborts the chain withUnhandledInferenceFinishReasoninstead of being retried or degraded gracefully. In practice this meansmax_output_tokensmust comfortably fit within the model’s own output limit, or a long response gets truncated and the session fails outright. -
Long context. The generator receives the composed conventions, the repository excerpts it reads, and the plan output; models with short context windows truncate silently.
-
Optional: native
web_search_optionssupport, which LiteLLM reports per model. Without it the web-search tool falls back to the OpenAI Responses API path, which only works on OpenAI-compatible endpoints.
The shipped models (Claude and Gemini on Vertex AI, vertex_ai/, which
the core’s test suite exercises) satisfy all of these, as do the built-in defaults (Anthropic Claude direct,
anthropic/) and the same Claude models on Azure AI Foundry
(azure_ai/). Kumoss has not tested every provider LiteLLM lists.
The provider/model-id notation
A LiteLLM model string is <provider>/<model-id>:
-
<provider>selects the LiteLLM integration and therefore the credential variables (for exampleanthropic,vertex_ai,azure,bedrock,openai,ollama). -
<model-id>is whatever that provider calls the model or deployment (claude-sonnet-5,gpt-4o, a deployment name, an endpoint name).
Routing is purely prefix based. Kumoss does not validate that the model id exists; a typo surfaces as an error from the provider on the first call. The reference list of model ids is models.litellm.ai; the provider list with credential keys is the LiteLLM provider docs.
Model selection versus credentials
Model selection lives in config.yaml; credentials live in
core/.env. They must belong to the same provider:
# config.yaml (fictitious example)
llm:
model: "anthropic/claude-sonnet-5"
small_model: "anthropic/claude-haiku-4-5"
temperature: 0.1
max_output_tokens: 32000
# core/.env (fictitious example)
ANTHROPIC_API_KEY=sk-example-not-a-real-key
Changing a model identifier to a different provider means changing the
credential variables too. config.yaml is baked into the core image,
so a model change requires docker compose build core; a credential
change in core/.env only requires restarting the core container.
Startup credential validation and its limits
At boot the core calls litellm.validate_environment(model=…) for
each configured model (or for each llm.model_list entry when that
key is set). If LiteLLM reports missing environment variables, the
core refuses to start with:
LLM credentials missing from environment: ANTHROPIC_API_KEY. Set the listed env vars or configure llm.model_list in config.yaml.
Treat this check as a convenience, not a guarantee:
-
Some providers have no validation mapping in LiteLLM. For
azure_ai,cohere_chat,databricks,watsonx,sambanova,anyscale, andsnowflakethe check reports nothing missing even with no credentials set; the failure appears on the first LLM call instead. The prefix matters, not the vendor:cohere/<model>is validated (COHERE_API_KEY),cohere_chat/<model>is not. -
Unknown prefixes pass silently. A string such as
nonexistent_provider/xor a baremodel-namewithout a provider is not rejected at boot. -
Vertex AI checks only the project and location variables. Boot always requires
VERTEXAI_PROJECTandVERTEXAI_LOCATIONto be set; the check is a plain environment lookup and does not consult Application Default Credentials or a Google Cloud SDK installation.VERTEXAI_CREDENTIALSis not checked at boot: when it is unset, LiteLLM falls back to Application Default Credentials (GOOGLE_APPLICATION_CREDENTIALS) on the first call, and a missing credential file fails there. -
Only presence is checked, never validity. A wrong or expired key passes validation and fails on the first call.
-
Boot validation always uses the provider’s default variable names, even for
llm.model_listentries that reference custom variables throughos.environ/…. See the note under Advanced routing.
The exact behaviour depends on the installed LiteLLM version (the core
pins litellm>=1.93.0).
Provider-specific credential variables
The tables below list the LiteLLM model string format, the environment
variables LiteLLM reads by default, and the litellm_params keys you
would use instead when writing an llm.model_list entry. The lists
come from LiteLLM’s documentation. Kumoss has not tested every
provider listed here; the checked-in config.yaml and the code
defaults both target Anthropic directly (anthropic/), and the core’s
test suite
(core/tests/shared/config/test_llm_config.py) exercises the
OpenAI-style and Vertex AI validation paths.
Direct providers (frontier labs)
| Provider | Model string | Default env vars | litellm_params keys |
Notes |
|---|---|---|---|---|
Anthropic |
|
|
|
e.g. |
OpenAI |
|
|
|
|
xAI |
|
|
|
|
Mistral AI |
|
|
|
|
Cohere |
|
|
|
Boot validation covers the |
Cloud platforms
| Provider | Model string | Default env vars | litellm_params keys |
Notes |
|---|---|---|---|---|
Google Vertex AI |
|
|
|
|
Google AI Studio (Gemini API) |
|
|
|
|
Azure OpenAI |
|
|
|
|
Azure AI Foundry |
|
|
|
No boot-time validation mapping in LiteLLM. |
AWS Bedrock |
|
|
|
Boot validation checks only the two key variables; the region is
still needed at call time. |
AWS SageMaker |
|
|
|
|
Cloudflare Workers AI |
|
|
|
Hosted inference services
| Provider | Model string | Default env vars | litellm_params keys |
|---|---|---|---|
Groq |
|
|
|
DeepSeek |
|
|
|
Together AI |
|
|
|
Fireworks AI |
|
|
|
OpenRouter |
|
|
|
Perplexity AI |
|
|
|
Cerebras |
|
|
|
SambaNova |
|
|
|
DeepInfra |
|
|
|
Anyscale |
|
|
|
Replicate |
|
|
|
AI21 |
|
|
|
Databricks |
|
|
|
IBM watsonx |
|
|
|
Snowflake Cortex |
|
|
|
OpenAI-compatible and self-hosted endpoints
| Engine | Model string | Default env vars | litellm_params keys |
Notes |
|---|---|---|---|---|
Ollama |
|
|
|
Default base |
vLLM, LocalAI, LM Studio, FastChat, TabbyAPI, and other OpenAI-compatible servers |
|
|
|
Point |
Hugging Face (serverless or dedicated endpoints, TGI) |
|
|
|
Minimal configuration example
A single direct provider, both roles on the same vendor, no router customisation. This is the checked-in setup, and the simplest working one.
# config.yaml
llm:
model: "anthropic/claude-sonnet-5"
small_model: "anthropic/claude-haiku-4-5"
temperature: 0.1
max_output_tokens: 32000
# core/.env (fictitious)
ANTHROPIC_API_KEY=sk-example-not-a-real-key
An equivalent setup on Azure AI Foundry, whose credentials LiteLLM cannot validate at boot (a missing value fails on the first call):
llm:
model: "azure_ai/claude-sonnet-4-5"
small_model: "azure_ai/claude-haiku-4-5"
AZURE_AI_API_KEY=example-not-a-real-key
AZURE_AI_API_BASE=https://demo-platform.services.ai.azure.com
Equivalent example on Google Vertex AI:
llm:
model: "vertex_ai/claude-sonnet-4-5"
small_model: "vertex_ai/gemini-3.7-flash"
VERTEXAI_PROJECT=demo-platform-project
VERTEXAI_LOCATION=europe-west1
VERTEXAI_CREDENTIALS={"type":"service_account","project_id":"demo-platform-project","private_key":"<service-account-private-key-pem>","client_email":"kumoss@demo-platform-project.iam.gserviceaccount.example.invalid"}
For how to paste the key, and what the file-path alternative needs, see the Vertex AI example in Environment variables.
Advanced routing with llm.model_list
llm.model_list follows the
LiteLLM Router format and is
passed to the Router verbatim. Use it only when you need something the
two plain model strings cannot express:
-
Load balancing. Several entries sharing one
model_nameare load balanced by the Router. Router-level fallbacks between differentmodel_name`s are not available: the core passes only `model_listto the Router and exposes nofallbacksorrouting_strategysetting. If the alias named byllm.modelorllm.small_modelhas no healthy deployment, the call fails. -
Custom credential variable names. A
litellm_paramsvalue of the formos.environ/VARIABLE_NAMEis resolved by LiteLLM from the environment at call time, so secrets stay out ofconfig.yaml. -
Per-model endpoints, such as a self-hosted OpenAI-compatible server for the small model and a cloud provider for the main model.
You do not need it to keep credentials in core/.env: plain model
strings already read the provider’s default variables.
Rules when model_list is non-empty:
-
llm.modelandllm.small_modelmust equal amodel_namein the list. They are router aliases now, not provider strings. -
Only
litellm_params.modelis required in an entry.api_key,api_base,api_version, and the other credential keys are optional: when omitted, LiteLLM reads the provider’s default environment variables at call time (forazure/…, that isAZURE_API_KEY,AZURE_API_BASE, andAZURE_API_VERSION). Add them only to point an entry at a different endpoint or at a differently named variable. -
Boot validation runs against each entry’s
litellm_params.modelinstead of the two role strings. -
Boot validation still checks the provider’s default variable names. If an entry references
os.environ/MY_CUSTOM_KEY, the default-named variable (for exampleANTHROPIC_API_KEY) must also be set, or the core will not boot. This is a LiteLLM limitation;validate_environmenttreats an empty-string value as present, so the default-named variable can hold any value, even empty, because only the custom one is used at call time.
If all you want is the default provider credentials from core/.env,
you do not need model_list at all; the two plain model strings are
the right configuration.
Aliases with default credentials
Two aliases and nothing else. Credentials come from core/.env under
the provider’s default names, exactly as without model_list:
# config.yaml (fictitious)
llm:
model: "main"
small_model: "small"
model_list:
- model_name: "main"
litellm_params:
model: "azure/main-deployment"
- model_name: "small"
litellm_params:
model: "azure/small-deployment"
# core/.env (fictitious)
AZURE_API_KEY=sk-example-not-a-real-key
AZURE_API_BASE=https://demo-platform.openai.azure.example.invalid
AZURE_API_VERSION=2025-01-01-preview
Advanced example: per-entry endpoints and custom variable names
Two aliases, two load-balanced Azure OpenAI deployments for the main
role in different regions, a self-hosted small model, and custom
credential variable names. Here api_key, api_base, and
api_version are present because each entry needs its own endpoint
and key; they would be redundant otherwise. All values are fictitious;
main-deployment is an Azure OpenAI deployment name.
# config.yaml (fictitious)
llm:
model: "main"
small_model: "small"
temperature: 0.1
max_output_tokens: 32000
model_list:
# Primary deployment for the main role.
- model_name: "main"
litellm_params:
model: "azure/main-deployment"
api_key: "os.environ/PLATFORM_AZURE_KEY_EU"
api_base: "https://eu.llm.example.invalid"
api_version: "2025-01-01-preview"
# Second deployment with the same alias: load balanced with the first.
- model_name: "main"
litellm_params:
model: "azure/main-deployment"
api_key: "os.environ/PLATFORM_AZURE_KEY_US"
api_base: "https://us.llm.example.invalid"
api_version: "2025-01-01-preview"
# Self-hosted OpenAI-compatible server for the small role.
- model_name: "small"
litellm_params:
model: "openai/llama-3.3-70b-instruct"
api_key: "os.environ/SELF_HOSTED_LLM_KEY"
api_base: "https://llm.example.invalid/v1"
# core/.env (fictitious)
PLATFORM_AZURE_KEY_EU=sk-example-not-a-real-key-eu
PLATFORM_AZURE_KEY_US=sk-example-not-a-real-key-us
SELF_HOSTED_LLM_KEY=sk-example-not-a-real-key-local
# Required only to satisfy boot validation, which checks the provider
# default names even when os.environ/... references are used.
AZURE_API_KEY=placeholder-for-boot-validation
AZURE_API_BASE=https://placeholder.example.invalid
AZURE_API_VERSION=2025-01-01-preview
OPENAI_API_KEY=placeholder-for-boot-validation
Troubleshooting
Core exits with LLM credentials missing from environment: …. The
provider chosen by llm.model or llm.small_model needs the listed
variables in core/.env. Either you changed the model prefix without
changing the credentials, or you rely on custom variable names through
model_list without also setting the default-named ones.
Core boots but the first session fails with an authentication or
"model not found" error. The provider has no boot-time validation
mapping (for example azure_ai), the key is present but invalid, or
the <model-id> part of the string is wrong for that provider. Check
the model id against models.litellm.ai and
the provider’s own console.
Boot validation complains about a default variable although you use
custom names. Expected. Set the default-named variable to any value,
even empty (validate_environment treats an empty string as present);
only the os.environ/… reference in model_list is used at call
time.
model or small_model does not match any model_name. When
model_list is set, the two role strings are aliases. LiteLLM raises
an error on the first call for an alias with no matching entry. Make
both values equal to a model_name from the list.
Requests time out or the endpoint is unreachable. Each chat
completion call has an llm.timeout budget (600 seconds by default)
and num_retries=3 inside LiteLLM. The web-search Responses API
fallback path uses the same timeout but is otherwise different: it
retries up to 4 attempts with 30-second sleeps between them and sets no
num_retries, so a slow or flaky OpenAI-compatible endpoint can stall
that path for a long time. For self-hosted endpoints, remember that the
core runs inside the compose network: localhost refers to the core
container, not to the host. Use the service name or the host’s
routable address in api_base, and make sure any TLS certificate is
trusted by the core image.
Provider rejects temperature or max_tokens. The core sends
drop_params=True, so unsupported parameters are dropped. If the
provider still rejects a value (for example a completion cap below
32000), lower llm.max_output_tokens.
Running outside Docker
litellm calls load_dotenv() at import time unless LITELLM_MODE
is set to PRODUCTION. Running the core, or pytest, on your
workstation (rather than inside the container) therefore silently
loads core/.env into the process environment — including any real
secrets it holds — even though nothing in Kumoss’s own code does this.
The working directory does not matter: python-dotenv searches upwards
from the directory of the calling package, so with the virtual
environment at core/.venv the file core/.env is found from any
directory. Variables already present in the environment are not
overwritten by the file. To reproduce a clean-environment boot, set
LITELLM_MODE=PRODUCTION, move core/.env aside, or use a virtual
environment outside core/; otherwise a variable you think is unset
may actually be coming from core/.env.