prerelease Prerelease stable Latest
Kumoss
prerelease Prerelease stable Latest

LLM providers and models

How Kumoss selects LLM models through LiteLLM, where credentials go, what the core checks at startup, and the llm.model_list router escape hatch.

Kumoss calls large language models (LLMs) through LiteLLM, a routing library that speaks to many providers behind one interface. This page explains how Kumoss selects models, where credentials go, what the core checks at startup, and how to use the advanced llm.model_list router configuration.

All values in the examples are fictitious. Domains use .invalid, and keys look like sk-example-not-a-real-key. Never paste real secrets into config.yaml or into documentation.

Related guides: Configuration (every config.yaml field), Environment variables and secrets (the .env files), and the two installation guides, Quickstart and Deploy to production, which say where the credentials live in each deployment model.

How Kumoss uses LiteLLM

Kumoss configures two model roles in config.yaml. Both are applied by the core’s LiteLLM Router; there is no provider enum in Kumoss, so any model string LiteLLM understands is accepted by the configuration. Whether a model works is a different matter; see Model requirements below.

Key Role Used for

llm.model

High-quality model

The IaC generator chain (writing and fixing Terraform code), the target generator chain (choosing plan targets), the report generator chain, and the task splitter chain (turning a drift report into remediation operations).

llm.small_model

Fast, cheaper model

Every other chain: request filtering, drift reconciliation filtering, drift exception filtering, prompt composition, pull-request text, the compliance checker, status messages, and the LLM calls that some tools make internally (for example web search).

Three extra settings apply to both roles:

Key Default Meaning

llm.temperature

0.1

Sampling temperature sent with every chat completion. When a chain requests extended thinking, the core sends 1.0 instead, because reasoning-enabled models require it, together with reasoning_effort: low. Tool-internal LLM calls use a fixed 0.5.

llm.max_output_tokens

32000

Sent as max_tokens (or max_output_tokens for the web-search path) on every call. Lower it if your provider caps completions below this value.

llm.timeout

600

Seconds LiteLLM waits for a single inference call before aborting it. Each of the three internal retries gets a fresh budget, so the worst case for one inference is four times this value. Raise it for slow reasoning models, lower it to fail faster on a stalled provider.

The core passes drop_params=True to LiteLLM, so a provider that does not support one of these parameters ignores it instead of failing.

Model requirements

Kumoss’s agents are tool-calling loops, so the configuration accepting a model string does not mean the model can run a session. Both roles must provide:

  • Reliable OpenAI-style function calling. Agent loops send a tools list with tool_choice: auto and expect well-formed JSON arguments back; each loop ends only when the model calls its sentinel tool. A reply without tool calls is answered with a reminder and costs a turn of the loop budget (see Agent tool loop), so a model that often answers in text instead of calling tools runs out of turns. This applies to agent loops specifically — the status/summary text calls (generate_text) send no tools at all. Models without reliable function calling (many small self-hosted models) fail every agent session.

  • A completion cap of at least llm.max_output_tokens (32000 by default), or lower that setting to what the provider allows.

  • A finish reason of tool_use or end_turn. Any other finish reason — length (the completion was truncated), content_filter, and so on — aborts the chain with UnhandledInferenceFinishReason instead of being retried or degraded gracefully. In practice this means max_output_tokens must comfortably fit within the model’s own output limit, or a long response gets truncated and the session fails outright.

  • Long context. The generator receives the composed conventions, the repository excerpts it reads, and the plan output; models with short context windows truncate silently.

  • Optional: native web_search_options support, which LiteLLM reports per model. Without it the web-search tool falls back to the OpenAI Responses API path, which only works on OpenAI-compatible endpoints.

The shipped models (Claude and Gemini on Vertex AI, vertex_ai/, which the core’s test suite exercises) satisfy all of these, as do the built-in defaults (Anthropic Claude direct, anthropic/) and the same Claude models on Azure AI Foundry (azure_ai/). Kumoss has not tested every provider LiteLLM lists.

The provider/model-id notation

A LiteLLM model string is <provider>/<model-id>:

  • <provider> selects the LiteLLM integration and therefore the credential variables (for example anthropic, vertex_ai, azure, bedrock, openai, ollama).

  • <model-id> is whatever that provider calls the model or deployment (claude-sonnet-5, gpt-4o, a deployment name, an endpoint name).

Routing is purely prefix based. Kumoss does not validate that the model id exists; a typo surfaces as an error from the provider on the first call. The reference list of model ids is models.litellm.ai; the provider list with credential keys is the LiteLLM provider docs.

Model selection versus credentials

Model selection lives in config.yaml; credentials live in core/.env. They must belong to the same provider:

# config.yaml (fictitious example)
llm:
  model: "anthropic/claude-sonnet-5"
  small_model: "anthropic/claude-haiku-4-5"
  temperature: 0.1
  max_output_tokens: 32000
# core/.env (fictitious example)
ANTHROPIC_API_KEY=sk-example-not-a-real-key

Changing a model identifier to a different provider means changing the credential variables too. config.yaml is baked into the core image, so a model change requires docker compose build core; a credential change in core/.env only requires restarting the core container.

Startup credential validation and its limits

At boot the core calls litellm.validate_environment(model=…​) for each configured model (or for each llm.model_list entry when that key is set). If LiteLLM reports missing environment variables, the core refuses to start with:

LLM credentials missing from environment: ANTHROPIC_API_KEY. Set the listed env vars or configure llm.model_list in config.yaml.

Treat this check as a convenience, not a guarantee:

  • Some providers have no validation mapping in LiteLLM. For azure_ai, cohere_chat, databricks, watsonx, sambanova, anyscale, and snowflake the check reports nothing missing even with no credentials set; the failure appears on the first LLM call instead. The prefix matters, not the vendor: cohere/<model> is validated (COHERE_API_KEY), cohere_chat/<model> is not.

  • Unknown prefixes pass silently. A string such as nonexistent_provider/x or a bare model-name without a provider is not rejected at boot.

  • Vertex AI checks only the project and location variables. Boot always requires VERTEXAI_PROJECT and VERTEXAI_LOCATION to be set; the check is a plain environment lookup and does not consult Application Default Credentials or a Google Cloud SDK installation. VERTEXAI_CREDENTIALS is not checked at boot: when it is unset, LiteLLM falls back to Application Default Credentials (GOOGLE_APPLICATION_CREDENTIALS) on the first call, and a missing credential file fails there.

  • Only presence is checked, never validity. A wrong or expired key passes validation and fails on the first call.

  • Boot validation always uses the provider’s default variable names, even for llm.model_list entries that reference custom variables through os.environ/…​. See the note under Advanced routing.

The exact behaviour depends on the installed LiteLLM version (the core pins litellm>=1.93.0).

Provider-specific credential variables

The tables below list the LiteLLM model string format, the environment variables LiteLLM reads by default, and the litellm_params keys you would use instead when writing an llm.model_list entry. The lists come from LiteLLM’s documentation. Kumoss has not tested every provider listed here; the checked-in config.yaml and the code defaults both target Anthropic directly (anthropic/), and the core’s test suite (core/tests/shared/config/test_llm_config.py) exercises the OpenAI-style and Vertex AI validation paths.

Direct providers (frontier labs)

Provider Model string Default env vars litellm_params keys Notes

Anthropic

anthropic/<model>

ANTHROPIC_API_KEY

api_key

e.g. anthropic/claude-sonnet-5. ANTHROPIC_AUTH_TOKEN is also accepted as an alternative to the API key.

OpenAI

openai/<model>

OPENAI_API_KEY, OPENAI_API_BASE

api_key, api_base

api_base defaults to https://api.openai.com/v1

xAI

xai/<model>

XAI_API_KEY

api_key

Mistral AI

mistral/<model>

MISTRAL_API_KEY

api_key

Cohere

cohere/<model> or cohere_chat/<model>

COHERE_API_KEY

api_key

Boot validation covers the cohere/ prefix only; cohere_chat/ has no validation mapping and a missing key fails on the first call.

Cloud platforms

Provider Model string Default env vars litellm_params keys Notes

Google Vertex AI

vertex_ai/<model>

VERTEXAI_PROJECT, VERTEXAI_LOCATION, VERTEXAI_CREDENTIALS

vertex_project, vertex_location, vertex_credentials

VERTEXAI_CREDENTIALS holds the service-account key JSON content, on one line. A path to a key file also works, but only if you mount that file into the core container yourself. Omit it to use Application Default Credentials (GOOGLE_APPLICATION_CREDENTIALS, workload identity, and so on).

Google AI Studio (Gemini API)

gemini/<model>

GEMINI_API_KEY (LiteLLM also accepts GOOGLE_API_KEY)

api_key

Azure OpenAI

azure/<deployment>

AZURE_API_KEY, AZURE_API_BASE, AZURE_API_VERSION

api_key, api_base, api_version

api_base is https://<resource>.openai.azure.com

Azure AI Foundry

azure_ai/<model>

AZURE_AI_API_KEY, AZURE_AI_API_BASE

api_key, api_base

No boot-time validation mapping in LiteLLM.

AWS Bedrock

bedrock/<model-id>

AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY, AWS_REGION_NAME

aws_access_key_id, aws_secret_access_key, aws_region_name

Boot validation checks only the two key variables; the region is still needed at call time. AWS_PROFILE, or AWS_ROLE_ARN
AWS_WEB_IDENTITY_TOKEN_FILE, are also accepted as alternatives to static keys.

AWS SageMaker

sagemaker/<endpoint>

AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY, AWS_REGION_NAME

aws_access_key_id, aws_secret_access_key, aws_region_name

Cloudflare Workers AI

cloudflare/<model>

CLOUDFLARE_API_KEY, CLOUDFLARE_API_BASE

api_key, api_base

Hosted inference services

Provider Model string Default env vars litellm_params keys

Groq

groq/<model>

GROQ_API_KEY

api_key

DeepSeek

deepseek/<model>

DEEPSEEK_API_KEY

api_key

Together AI

together_ai/<model>

TOGETHERAI_API_KEY

api_key

Fireworks AI

fireworks_ai/<model>

FIREWORKS_AI_API_KEY

api_key

OpenRouter

openrouter/<model>

OPENROUTER_API_KEY

api_key

Perplexity AI

perplexity/<model>

PERPLEXITYAI_API_KEY

api_key

Cerebras

cerebras/<model>

CEREBRAS_API_KEY

api_key

SambaNova

sambanova/<model>

SAMBANOVA_API_KEY

api_key

DeepInfra

deepinfra/<model>

DEEPINFRA_API_KEY

api_key

Anyscale

anyscale/<model>

ANYSCALE_API_KEY

api_key

Replicate

replicate/<model>

REPLICATE_API_KEY

api_key

AI21

ai21/<model>

AI21_API_KEY

api_key

Databricks

databricks/<endpoint>

DATABRICKS_API_KEY, DATABRICKS_API_BASE

api_key, api_base

IBM watsonx

watsonx/<model-id>

WATSONX_APIKEY, WATSONX_URL, WATSONX_PROJECT_ID

api_key, api_base, watsonx_project_id

Snowflake Cortex

snowflake/<model>

SNOWFLAKE_ACCOUNT_ID, SNOWFLAKE_USER, SNOWFLAKE_PASSWORD

snowflake_account_id, snowflake_user, snowflake_password

OpenAI-compatible and self-hosted endpoints

Engine Model string Default env vars litellm_params keys Notes

Ollama

ollama/<model>

OLLAMA_API_BASE

api_base

Default base http://localhost:11434; no API key. From inside the compose network, use the host or container address, not localhost.

vLLM, LocalAI, LM Studio, FastChat, TabbyAPI, and other OpenAI-compatible servers

openai/<model>

OPENAI_API_KEY, OPENAI_API_BASE

api_key, api_base

Point api_base at http://<host>:<port>/v1. Boot validation still requires OPENAI_API_KEY to be set; use a placeholder such as none if the server does not check it.

Hugging Face (serverless or dedicated endpoints, TGI)

huggingface/<repo/model>

HUGGINGFACE_API_KEY

api_key, api_base

Minimal configuration example

A single direct provider, both roles on the same vendor, no router customisation. This is the checked-in setup, and the simplest working one.

# config.yaml
llm:
  model: "anthropic/claude-sonnet-5"
  small_model: "anthropic/claude-haiku-4-5"
  temperature: 0.1
  max_output_tokens: 32000
# core/.env (fictitious)
ANTHROPIC_API_KEY=sk-example-not-a-real-key

An equivalent setup on Azure AI Foundry, whose credentials LiteLLM cannot validate at boot (a missing value fails on the first call):

llm:
  model: "azure_ai/claude-sonnet-4-5"
  small_model: "azure_ai/claude-haiku-4-5"
AZURE_AI_API_KEY=example-not-a-real-key
AZURE_AI_API_BASE=https://demo-platform.services.ai.azure.com

Equivalent example on Google Vertex AI:

llm:
  model: "vertex_ai/claude-sonnet-4-5"
  small_model: "vertex_ai/gemini-3.7-flash"
VERTEXAI_PROJECT=demo-platform-project
VERTEXAI_LOCATION=europe-west1
VERTEXAI_CREDENTIALS={"type":"service_account","project_id":"demo-platform-project","private_key":"<service-account-private-key-pem>","client_email":"kumoss@demo-platform-project.iam.gserviceaccount.example.invalid"}

For how to paste the key, and what the file-path alternative needs, see the Vertex AI example in Environment variables.

Advanced routing with llm.model_list

llm.model_list follows the LiteLLM Router format and is passed to the Router verbatim. Use it only when you need something the two plain model strings cannot express:

  • Load balancing. Several entries sharing one model_name are load balanced by the Router. Router-level fallbacks between different model_name`s are not available: the core passes only `model_list to the Router and exposes no fallbacks or routing_strategy setting. If the alias named by llm.model or llm.small_model has no healthy deployment, the call fails.

  • Custom credential variable names. A litellm_params value of the form os.environ/VARIABLE_NAME is resolved by LiteLLM from the environment at call time, so secrets stay out of config.yaml.

  • Per-model endpoints, such as a self-hosted OpenAI-compatible server for the small model and a cloud provider for the main model.

You do not need it to keep credentials in core/.env: plain model strings already read the provider’s default variables.

Rules when model_list is non-empty:

  1. llm.model and llm.small_model must equal a model_name in the list. They are router aliases now, not provider strings.

  2. Only litellm_params.model is required in an entry. api_key, api_base, api_version, and the other credential keys are optional: when omitted, LiteLLM reads the provider’s default environment variables at call time (for azure/…​, that is AZURE_API_KEY, AZURE_API_BASE, and AZURE_API_VERSION). Add them only to point an entry at a different endpoint or at a differently named variable.

  3. Boot validation runs against each entry’s litellm_params.model instead of the two role strings.

  4. Boot validation still checks the provider’s default variable names. If an entry references os.environ/MY_CUSTOM_KEY, the default-named variable (for example ANTHROPIC_API_KEY) must also be set, or the core will not boot. This is a LiteLLM limitation; validate_environment treats an empty-string value as present, so the default-named variable can hold any value, even empty, because only the custom one is used at call time.

If all you want is the default provider credentials from core/.env, you do not need model_list at all; the two plain model strings are the right configuration.

Aliases with default credentials

Two aliases and nothing else. Credentials come from core/.env under the provider’s default names, exactly as without model_list:

# config.yaml (fictitious)
llm:
  model: "main"
  small_model: "small"
  model_list:
    - model_name: "main"
      litellm_params:
        model: "azure/main-deployment"
    - model_name: "small"
      litellm_params:
        model: "azure/small-deployment"
# core/.env (fictitious)
AZURE_API_KEY=sk-example-not-a-real-key
AZURE_API_BASE=https://demo-platform.openai.azure.example.invalid
AZURE_API_VERSION=2025-01-01-preview

Advanced example: per-entry endpoints and custom variable names

Two aliases, two load-balanced Azure OpenAI deployments for the main role in different regions, a self-hosted small model, and custom credential variable names. Here api_key, api_base, and api_version are present because each entry needs its own endpoint and key; they would be redundant otherwise. All values are fictitious; main-deployment is an Azure OpenAI deployment name.

# config.yaml (fictitious)
llm:
  model: "main"
  small_model: "small"
  temperature: 0.1
  max_output_tokens: 32000
  model_list:
    # Primary deployment for the main role.
    - model_name: "main"
      litellm_params:
        model: "azure/main-deployment"
        api_key: "os.environ/PLATFORM_AZURE_KEY_EU"
        api_base: "https://eu.llm.example.invalid"
        api_version: "2025-01-01-preview"
    # Second deployment with the same alias: load balanced with the first.
    - model_name: "main"
      litellm_params:
        model: "azure/main-deployment"
        api_key: "os.environ/PLATFORM_AZURE_KEY_US"
        api_base: "https://us.llm.example.invalid"
        api_version: "2025-01-01-preview"
    # Self-hosted OpenAI-compatible server for the small role.
    - model_name: "small"
      litellm_params:
        model: "openai/llama-3.3-70b-instruct"
        api_key: "os.environ/SELF_HOSTED_LLM_KEY"
        api_base: "https://llm.example.invalid/v1"
# core/.env (fictitious)
PLATFORM_AZURE_KEY_EU=sk-example-not-a-real-key-eu
PLATFORM_AZURE_KEY_US=sk-example-not-a-real-key-us
SELF_HOSTED_LLM_KEY=sk-example-not-a-real-key-local

# Required only to satisfy boot validation, which checks the provider
# default names even when os.environ/... references are used.
AZURE_API_KEY=placeholder-for-boot-validation
AZURE_API_BASE=https://placeholder.example.invalid
AZURE_API_VERSION=2025-01-01-preview
OPENAI_API_KEY=placeholder-for-boot-validation

Troubleshooting

Core exits with LLM credentials missing from environment: …​. The provider chosen by llm.model or llm.small_model needs the listed variables in core/.env. Either you changed the model prefix without changing the credentials, or you rely on custom variable names through model_list without also setting the default-named ones.

Core boots but the first session fails with an authentication or "model not found" error. The provider has no boot-time validation mapping (for example azure_ai), the key is present but invalid, or the <model-id> part of the string is wrong for that provider. Check the model id against models.litellm.ai and the provider’s own console.

Boot validation complains about a default variable although you use custom names. Expected. Set the default-named variable to any value, even empty (validate_environment treats an empty string as present); only the os.environ/…​ reference in model_list is used at call time.

model or small_model does not match any model_name. When model_list is set, the two role strings are aliases. LiteLLM raises an error on the first call for an alias with no matching entry. Make both values equal to a model_name from the list.

Requests time out or the endpoint is unreachable. Each chat completion call has an llm.timeout budget (600 seconds by default) and num_retries=3 inside LiteLLM. The web-search Responses API fallback path uses the same timeout but is otherwise different: it retries up to 4 attempts with 30-second sleeps between them and sets no num_retries, so a slow or flaky OpenAI-compatible endpoint can stall that path for a long time. For self-hosted endpoints, remember that the core runs inside the compose network: localhost refers to the core container, not to the host. Use the service name or the host’s routable address in api_base, and make sure any TLS certificate is trusted by the core image.

Provider rejects temperature or max_tokens. The core sends drop_params=True, so unsupported parameters are dropped. If the provider still rejects a value (for example a completion cap below 32000), lower llm.max_output_tokens.

Running outside Docker

litellm calls load_dotenv() at import time unless LITELLM_MODE is set to PRODUCTION. Running the core, or pytest, on your workstation (rather than inside the container) therefore silently loads core/.env into the process environment — including any real secrets it holds — even though nothing in Kumoss’s own code does this. The working directory does not matter: python-dotenv searches upwards from the directory of the calling package, so with the virtual environment at core/.venv the file core/.env is found from any directory. Variables already present in the environment are not overwritten by the file. To reproduce a clean-environment boot, set LITELLM_MODE=PRODUCTION, move core/.env aside, or use a virtual environment outside core/; otherwise a variable you think is unset may actually be coming from core/.env.