C CORTANADocumentation › Guides

Configuration

Cortana runs with no configuration at all. When you want to change something, use cortana.yml for structured settings and .env for secrets and quick overrides.

Where settings come from

Each source overrides the one before it:

  1. Built-in defaults
  2. cortana.yml (or cortana.yaml) in the workspace, the project's copy, --config PATH or CORTANA_CONFIG
  3. .env and the real environment
  4. Command-line flags such as --model or --memory, and live changes with /config, /model or /permissions

Start from the committed template:

cp cortana.example.yml cortana.yml

Keep secrets such as QDRANT_API_KEY, TTS_HTTP_API_KEY or MCP tokens in .env. cortana.yml, .env and .cortana/ are git-ignored; don't force-add them.

Model & reasoning

The default chat model is gpt-oss:20b. Change it with agent.model, OLLAMA_MODEL, --model, or /model during a session.

ModelSettings and Reasoning mirror the OpenAI Agents SDK. Each engine maps those controls to the parameters its API supports:

from libs import Agent, ModelSettings, Reasoning

agent = Agent(
    name="deep-researcher",
    model="gpt-oss:20b",
    model_settings=ModelSettings(reasoning=Reasoning(mode="pro", effort="medium", context="current_turn")),
)
  • effort: minimal becomes low; xhigh and max become high.
  • context="current_turn" (default) keeps reasoning between a turn's tool calls and drops it afterwards, which is how models such as gpt-oss are trained. all_turns replays every old trace and tends to cause overthinking.

On gpt-oss:20b, medium effort was much faster than high/max without producing worse plans.

During a session:

CommandEffect
/config thinking on|off|autoon is medium effort; auto starts low and lets the model raise it with set_reasoning_effort
/config reasoning none|minimal|low|medium|high|xhigh|maxA fixed level for later turns and new subagents
/config show-thinking on|offShow or hide reasoning in the transcript only

Docker Model Runner

Select Docker Model Runner as the chat and embedding engine. Docker Desktop users must enable Model Runner and host TCP access on port 12434. Pull the models you want to use:

docker model pull ai/smollm2
docker model pull ai/nomic-embed-text-v2-moe

Configure Cortana with Docker Hub model IDs, including their namespace:

agent:
  engine: docker-model-runner
  model: ai/smollm2
  docker_model_runner:
    url: http://localhost:12434

memory:
  provider: qdrant
  semantic_recall:
    enabled: true
    embed_model: ai/nomic-embed-text-v2-moe
  qdrant:
    collection: cortana_memory_dmr

The default URL is for Cortana running on the host. A Docker Desktop container can use http://model-runner.docker.internal. Override the URL with DOCKER_MODEL_RUNNER_URL.

The Nomic embedding model returns 768-dimensional vectors. Use a new Qdrant collection or re-embed existing memories when changing from a 1024-dimensional model such as mxbai-embed-large; vectors from different embedding models cannot be mixed.

The DMR OpenAI-compatible API supports chat, tool calls, streaming and embeddings. Its documented parameter set is narrower than llama.cpp: strict JSON-schema and model-specific reasoning controls are not available. The example models above do not support vision; image analysis needs a vision-capable DMR model.

Ollama setup

With semantic recall on, Cortana uses two Ollama runners: one pinned for embeddings and one rotating between chat, vision and image models. Ollama's server limits must be set before it starts:

launchctl setenv OLLAMA_MAX_LOADED_MODELS 2
launchctl setenv OLLAMA_NUM_PARALLEL 1
osascript -e 'quit app "Ollama"'
open -a Ollama
agent:
  ollama:
    keep_alive: 10m          # chat model
    embed_keep_alive: -1     # keep the embedding model loaded
    max_loaded_models: 2
    num_parallel: 1

The first request after a model swap (for example to the vision model) is slower while the replacement loads.

Tell Cortana the context size Ollama actually uses so it can show usage and compact in time:

OLLAMA_CONTEXT_WINDOW=32768 cortana

Point at another server with OLLAMA_HOST. /doctor checks the connection and whether the model supports tools and thinking.

llama.cpp

agent.engine: llamacpp runs models with llama-server instead of Ollama. Cortana starts the server with the app and stops it on exit, serving one GGUF model to every role. Install it with brew install llama.cpp, then:

agent:
  model: gemma4:26b-mlx      # still the name /model and /doctor show
  context_window: 32768
  engine: llamacpp
  llamacpp:
    hf_repo: ggml-org/gemma-3-4b-it-GGUF:Q4_K_M
    parallel: 4

On first launch Cortana downloads the model into model/llamacpp/<org>/<repo>/, with progress in the transcript, then starts the server. Later launches use the file on disk. The UI starts right away; the first model call waits until the model has loaded.

KeyWhat it does
hf_repo<org>/<repo>[:<quant>], llama.cpp's -hf spelling. Without a quant, Q4_K_M is preferred.
hf_fileOne exact file in the repository instead of a quant
model_pathA GGUF you already have; skips downloading
visionAlso download the vision projector (mmproj) that view_image needs
parallelSlots that generate at once, sharing one KV cache. Keep it near agent.max_parallel_agents.
cache_type_k / cache_type_vq8_0 roughly halves the KV cache
urlUse a llama-server you run yourself; Cortana starts none

Set HF_TOKEN in .env for gated repositories. The server log is .cortana/llama-server-<port>.log, and a server that dies is restarted on the next request.

Chat templates

llama-server renders the Jinja template inside the GGUF, and some are strict. Gemma 3's, for example, allows one leading system message and strictly alternating turns. llamacpp.chat_template.messages controls how Cortana's history is converted:

  • native sends system, user, assistant and tool roles as is. Qwen, gpt-oss and Llama 3.x take it.
  • alternating merges system messages and turns tool results into user turns, for Gemma 3 and older Mistral templates.
  • auto (default) starts native and switches to alternating when the template rejects the turn order.

One model answers every role: helper, verifier and vision models named in the config use the loaded GGUF. Image generation stays on Ollama or Qwen-Image, and embeddings stay on Ollama unless llamacpp.embedding sets a model.

cortana.yml reference

The top-level sections. cortana.example.yml documents every key with comments.

SectionControlsDocs
agentModel, engine, reasoning, parallelism, tool iterations, permission mode, destructive-command confirmation, tool builder, context window, extra dirs, planning and verifierCore, Planning
memoryProvider, history length, working memory, semantic recall, chunking, QdrantMemory
voiceTTS provider and mode, Kokoro/Qwen voices, HTTP server, hands-free listeningVoice
imageProvider, size, steps, device, output folder, vision modelImages
videoWan model, size, frames, fps, stepsVideos
heartbeatCheck-in interval and active hoursHeartbeat
learningExperience recipesLearning
uiShow thinking, recaps, tips, helper model, /clear hintCLI
graphifyKnowledge graph, context mode, tools, excludesKnowledge graph
browserEngine, channel, headless, idle timeoutBrowser
mcpEnable, connect timeout, serversMCP

A typical starting point:

agent:
  model: gpt-oss:20b
  reasoning:
    effort: medium
    context: current_turn
  max_parallel_tools: 6
  max_tool_iterations: 10
  permission_mode: auto      # auto | plan | ask
  confirm_destructive: true
  context_window: 32768
  planning: true
  # extra_dirs: [~/Documents/notes]

memory:
  provider: local            # none | local | qdrant
  semantic_recall:
    enabled: true
    embed_model: mxbai-embed-large:335m

voice:
  provider: kokoro

ui:
  show_thinking: true
  recaps: true
  tips: true

Environment variables

Every setting has an environment override, which wins over cortana.yml. The most useful:

VariablePurpose
OLLAMA_MODELChat model
OLLAMA_HOSTOllama server address
CORTANA_ENGINEollama, docker-model-runner or llamacpp
DOCKER_MODEL_RUNNER_URLDocker Model Runner host URL
OLLAMA_CONTEXT_WINDOWEffective context size, for usage display and compaction
OLLAMA_THINKTurn model reasoning on or off
OLLAMA_KEEP_ALIVEHow long the chat model stays loaded
CONTEXT_COMPACTION_THRESHOLDCompaction trigger, 0–1 (default 0.9)
REASONING_EFFORT, REASONING_CONTEXTReasoning defaults
PERMISSION_MODEauto, plan or ask
CONFIRM_DESTRUCTIVEAsk before destructive commands
CORTANA_CONFIGPath to the YAML settings file
CORTANA_EXTRA_DIRSMore folders the file tools may use
CORTANA_STATE_DIRWhere schedules and background tasks are stored
MEMORY_PROVIDERnone, local or qdrant
MEMORY_EMBED_MODELEmbedding model for recall
QDRANT_URL, QDRANT_API_KEYQdrant server and credential
TTS_PROVIDER, TTS_MODEVoice provider (kokoro/qwen) and local/http
KOKORO_TTS_VOICEKokoro voice
IMAGE_PROVIDER, IMAGE_VISION_MODELImage provider and vision model
BROWSER_ENABLED, BROWSER_HEADLESSBrowser tools
HEARTBEAT_ENABLEDSelf-started check-ins
MCP_ENABLEDMCP servers

A TTS_PROVIDER=qwen left in .env or your shell profile wins over voice.provider in cortana.yml. Remove it to get Kokoro back.

Running through Temporal

For a dashboard of workflow status and worker progress, run turns through Temporal. Start the server and UI with Docker, then a worker:

docker compose -f docker-compose.temporal.yml up -d
uv run --extra temporal python -m app.temporal_runner worker

Submit a normal turn from another terminal:

uv run --extra temporal python -m app.temporal_runner run \
  --workspace "$PWD" "Explain how this project is structured"

Open http://localhost:8080 to see the durable workflow and activity history. The worker prints structured events (run_started, tool_started, run_finished, …). Set TEMPORAL_ADDRESS when the server runs on another machine.