Skip to content

Setting Up the Language Model

File: configs/llm.yaml
Command: kenzy-llm [config_path]

The language model is Kenzy's "thinking" — and whose computer it thinks on is your call. Kenzy speaks LiteLLM, so any provider works, including one running entirely in your house.

Cloud or local?

Cloud (OpenAI, Anthropic, OpenRouter) Local (Ollama, LM Studio)
What it needs An API key. No downloads. A GPU with 8–16 GB VRAM, or an Apple-silicon Mac. See the hardware note below.
What leaves your network The transcribed text of each request. With the default local speech recognition, your audio never leaves home. Nothing.
What it costs Per-token billing from the provider. Electricity and the hardware you already own.

The default is cloud — one key, no downloads, and it runs beside everything else on a small box. That makes it the fastest route to a working system, not a recommendation about where your data belongs.

Local models and tool calling

This is the one service where local costs real iron. Kenzy's skills depend on tool calling, and small models are noticeably worse at it. A ~7–14B model on a GPU gives a good experience; CPU-only LLM inference is technically possible and practically frustrating. (Speech recognition and the voice, by contrast, are perfectly happy on a CPU.)

Deciding this across all the services at once? Start at Running Fully Local.

Set it up

Set model to a provider's model string and give Kenzy that provider's API key. No base_url — LiteLLM infers the provider from the prefix.

model: "gpt-5-mini"        # or "claude-sonnet-5", "openrouter/…"

Requires: the matching key in ~/.config/kenzy/.envOPENAI_API_KEY, ANTHROPIC_API_KEY, or the provider's equivalent. Set it from the dashboard under Settings → API keys.

API keys, if you're new to them

An API key is a password that lets Kenzy use an account you hold with a provider. You create it in the provider's account pages (OpenAI: platform.openai.com → API keys; Anthropic: console.anthropic.com; OpenRouter: openrouter.ai → Keys), copy it once, and paste it into Kenzy's dashboard. Kenzy stores it on your server, never displays it again, and never sends it anywhere except to that provider.

Run Ollama or LM Studio, then point Kenzy at it:

curl -fsSL https://ollama.com/install.sh | sh
ollama pull qwen2.5:14b        # any tool-calling-capable model
model: "ollama/qwen2.5:14b"
base_url: "http://127.0.0.1:11434"

Requires: no API key. Skill sub-calls (news summaries, the Home Assistant resolver) follow the same model unless you override their per-skill model.

The end-to-end walkthrough, including picking a model, is Running Fully Local.

A cloud primary with a local fallback keeps Kenzy working through an internet outage. When the primary call fails, the request is silently retried once against the fallback, and the rest of that request stays there:

model: "gpt-5-mini"
fallback:
  model: "ollama/qwen2.5:14b"
  base_url: "http://127.0.0.1:11434"

See the hybrid option.

Custom endpoints never receive a cloud provider's API key

base_url redirects model calls to a server you choose. Requests to it deliberately carry none of the provider keys from your environment (OPENAI_API_KEY etc.) — so a changed or mistyped endpoint can never leak a cloud credential. Local providers (Ollama, LM Studio) need no key and just work. If your custom endpoint is a hosted proxy that requires auth (a LiteLLM proxy, OpenRouter), put its key in CUSTOM_LLM_API_KEY in .env — that is the one credential ever sent to a base_url. (If you previously kept such a proxy key in OPENAI_API_KEY, move it there.)

Pulled from the server

kenzy-llm pulls this config from the server at boot — it discovers the server via mDNS (or KENZY_SERVER_URL) and blocks until it answers, so start the server first. Edit it from the dashboard's Services tab (writes configs/services/llm.yaml on the server and restarts the service). Passing an explicit path loads locally instead (dev/offline). See central config for backend services.

Model strings

LiteLLM encodes the provider in the model string prefix:

Provider Example model string
OpenAI gpt-4o, gpt-4o-mini
Anthropic claude-opus-4-8, claude-sonnet-4-6
Ollama (local) ollama/hermes3, ollama/llama3.1
LM Studio (local) openai/model-name + base_url
Together AI together_ai/meta-llama/Llama-3-70b

API keys are read automatically from the environment. See the LiteLLM provider docs for the required environment variable per provider.

A complete example

host: "127.0.0.1"
port: 8766

model: "gpt-4o"
max_tool_iterations: 10

system_prompt: |
  You are Kenzy, a helpful home assistant. Be concise and conversational.

voice_prompt: "Speak in a friendly, natural tone at a moderate pace."

location:
  city: "Raleigh"
  state: "NC"
  country: "US"
  timezone: "America/New_York"
  latitude: 35.7796
  longitude: -78.6382

skills:
  dir: skills
  disabled: []

  weather:
    units: imperial

  home_assistant:
    url: "http://homeassistant.local:8123"
    curation_file: "data/home_assistant/curation.yaml"   # optional
    default_room: "living_room"

Advanced

Everything below is the full reference — useful when you're tuning; skippable when you're starting.

Full reference

Core

Key Default Description
host "127.0.0.1" Bind address
port 8766 HTTP port
log_level "info" What the service prints to its console
log_capture_level "debug" How deep the dashboard log viewer can see, independent of log_level
model "gpt-5.1" (shipped config) LiteLLM model string (see Model strings)
base_url Provider base URL. Required for Ollama, LM Studio, and similar local providers.
fallback.model Optional local fallback: when the primary model call fails (cloud outage, provider error), the request is silently retried once against this model — e.g. "ollama/qwen2.5:14b". If the fallback also fails, the user just hears the error cue. Unset = no fallback.
fallback.base_url The fallback model's endpoint, e.g. "http://127.0.0.1:11434"
params.reasoning_effort "" How long the model may "think" before speaking. Empty = don't send the parameter (models with adaptive defaults, like gpt-5.1, gain nothing from an explicit value). Set nonehigh to force a level on models whose default reasoning is heavier. Ignored harmlessly by providers that don't support it.
params.service_tier "" OpenAI service tier — "priority" is the paid low-latency tier if your account has it. Empty = don't send.
params.* Anything else LiteLLM accepts (service_tier: "priority" for OpenAI's paid low-latency tier, temperature, max_tokens, …), merged into every model call. Unsupported parameters are dropped per-provider. Credential/routing keys (api_key, base_url, …) are ignored here by design.
max_tool_iterations 10 Maximum skill call iterations per request before returning whatever the model has

Prompts

Key Default Description
system_prompt (built-in) Injected at the start of every LLM call. Defines Kenzy's persona and behavior.
voice_prompt (built-in) Fallback TTS style instruction used when the model does not provide one.

Location context

Injected into every LLM call and used as the default by location-aware skills (e.g. weather).

Key Description
location.city City name
location.state State or province
location.country Country code (e.g. "US")
location.timezone IANA timezone (e.g. "America/New_York")
location.latitude Decimal latitude
location.longitude Decimal longitude

Skills

Key Default Description
skills.dir "skills" Your skills overlay directory, loaded in addition to the bundled built-in skills. Relative paths resolve under the config home — ~/.config/kenzy/skills, or the repo root in a dev checkout.
skills.disabled [] Skill function names to disable (applies to built-in and overlay skills alike) without deleting any file.

Per-skill configuration lives under skills.<skill_name> as a nested map. See Built-in Skills for the keys each skill accepts.

Memory

Per-person memory — the fact ledger and its upkeep jobs (job history is visible on the service's token-gated GET /jobs).

Key Default Description
memory.enabled true The whole memory feature. false ⇒ no ledger, no memory skills, the /memory endpoints answer 503, and the dashboard's memory surfaces say so.
memory.file "data/memory/facts.jsonl" The ledger file, config-home-relative — plain JSONL, human-readable, rides backups.
memory.maintenance_interval 3600 Backstop cadence for the mechanical sweep (expired facts, exact duplicates, superseded tombstones past their keep window). No model involved. Each write already kicks the sweep, so this interval only has to catch what a write can't: a fact reaching its expiry, or a tombstone ageing out. 0 disables. (Was 60 s before 5.0.5 — if your config carries an explicit value, clear the key to follow the default.)
memory.superseded_keep_days 30 How long a superseded fact stays on disk (recoverable) before the sweep removes it.
memory.semantic_interval 86400 The daily backstop for semantic consolidation (merging restatements with your configured model). The real trigger is each "remember…" — this catches anything a failed run left behind. 0 disables the semantic layer entirely.
memory.semantic_cooldown 30 Rate limit between model-driven consolidation runs — dictating five facts in a row costs one model call, not five.
memory.private_to_cloud false By default, private-tier facts are withheld from a cloud model's context and consolidation (they still answer by voice, and consolidate on a local model). true opts out of the protection.
memory.classifier_model (blank) The write-path classifier's model — it judges whether a new memory contains a secret (release / vault to the lockbox / split). Blank = the service model, but a cloud model is never consulted for this: without a local model, ambiguous writes are held for dashboard review instead. Point this at a local model (e.g. ollama/qwen3:8b) to resolve ambiguity automatically.
memory.classifier_url (blank) The classifier model's endpoint (when classifier_model is set).
memory.classifier_keep_alive (blank) Ollama keep_alive sent with classifier/consolidation calls — "-1" pins the model resident, "30m" for half an hour. Blank = the Ollama server's own default (typically 5 minutes, meaning cold-start latency on sparse memory writes). Only sent to ollama/* models.