Configuration
Cortana runs with no configuration at all. When you want to change something, use cortana.yml for structured settings and .env for secrets and quick overrides.
Where settings come from
Each source overrides the one before it:
- Built-in defaults
cortana.yml(orcortana.yaml) in the workspace, the project's copy,--config PATHorCORTANA_CONFIG.envand the real environment- Command-line flags such as
--modelor--memory, and live changes with/config,/modelor/permissions
Start from the committed template:
cp cortana.example.yml cortana.yml
Keep secrets such as QDRANT_API_KEY, TTS_HTTP_API_KEY or MCP tokens in .env. cortana.yml, .env and .cortana/ are git-ignored; don't force-add them.
Model & reasoning
The default chat model is gpt-oss:20b. Change it with agent.model, OLLAMA_MODEL, --model, or /model during a session.
ModelSettings and Reasoning mirror the OpenAI Agents SDK. Each engine maps those controls to the parameters its API supports:
from libs import Agent, ModelSettings, Reasoning
agent = Agent(
name="deep-researcher",
model="gpt-oss:20b",
model_settings=ModelSettings(reasoning=Reasoning(mode="pro", effort="medium", context="current_turn")),
)
effort:minimalbecomeslow;xhighandmaxbecomehigh.context="current_turn"(default) keeps reasoning between a turn's tool calls and drops it afterwards, which is how models such as gpt-oss are trained.all_turnsreplays every old trace and tends to cause overthinking.
On gpt-oss:20b, medium effort was much faster than high/max without producing worse plans.
During a session:
| Command | Effect |
|---|---|
/config thinking on|off|auto | on is medium effort; auto starts low and lets the model raise it with set_reasoning_effort |
/config reasoning none|minimal|low|medium|high|xhigh|max | A fixed level for later turns and new subagents |
/config show-thinking on|off | Show or hide reasoning in the transcript only |
Docker Model Runner
Select Docker Model Runner as the chat and embedding engine. Docker Desktop users must enable Model Runner and host TCP access on port 12434. Pull the models you want to use:
docker model pull ai/smollm2
docker model pull ai/nomic-embed-text-v2-moe
Configure Cortana with Docker Hub model IDs, including their namespace:
agent:
engine: docker-model-runner
model: ai/smollm2
docker_model_runner:
url: http://localhost:12434
memory:
provider: qdrant
semantic_recall:
enabled: true
embed_model: ai/nomic-embed-text-v2-moe
qdrant:
collection: cortana_memory_dmr
The default URL is for Cortana running on the host. A Docker Desktop container can use http://model-runner.docker.internal. Override the URL with DOCKER_MODEL_RUNNER_URL.
The Nomic embedding model returns 768-dimensional vectors. Use a new Qdrant collection or re-embed existing memories when changing from a 1024-dimensional model such as mxbai-embed-large; vectors from different embedding models cannot be mixed.
The DMR OpenAI-compatible API supports chat, tool calls, streaming and embeddings. Its documented parameter set is narrower than llama.cpp: strict JSON-schema and model-specific reasoning controls are not available. The example models above do not support vision; image analysis needs a vision-capable DMR model.
Ollama setup
With semantic recall on, Cortana uses two Ollama runners: one pinned for embeddings and one rotating between chat, vision and image models. Ollama's server limits must be set before it starts:
launchctl setenv OLLAMA_MAX_LOADED_MODELS 2
launchctl setenv OLLAMA_NUM_PARALLEL 1
osascript -e 'quit app "Ollama"'
open -a Ollama
agent:
ollama:
keep_alive: 10m # chat model
embed_keep_alive: -1 # keep the embedding model loaded
max_loaded_models: 2
num_parallel: 1
The first request after a model swap (for example to the vision model) is slower while the replacement loads.
Tell Cortana the context size Ollama actually uses so it can show usage and compact in time:
OLLAMA_CONTEXT_WINDOW=32768 cortana
Point at another server with OLLAMA_HOST. /doctor checks the connection and whether the model supports tools and thinking.
llama.cpp
agent.engine: llamacpp runs models with llama-server instead of Ollama. Cortana starts the server with the app and stops it on exit, serving one GGUF model to every role. Install it with brew install llama.cpp, then:
agent:
model: gemma4:26b-mlx # still the name /model and /doctor show
context_window: 32768
engine: llamacpp
llamacpp:
hf_repo: ggml-org/gemma-3-4b-it-GGUF:Q4_K_M
parallel: 4
On first launch Cortana downloads the model into model/llamacpp/<org>/<repo>/, with progress in the transcript, then starts the server. Later launches use the file on disk. The UI starts right away; the first model call waits until the model has loaded.
| Key | What it does |
|---|---|
hf_repo | <org>/<repo>[:<quant>], llama.cpp's -hf spelling. Without a quant, Q4_K_M is preferred. |
hf_file | One exact file in the repository instead of a quant |
model_path | A GGUF you already have; skips downloading |
vision | Also download the vision projector (mmproj) that view_image needs |
parallel | Slots that generate at once, sharing one KV cache. Keep it near agent.max_parallel_agents. |
cache_type_k / cache_type_v | q8_0 roughly halves the KV cache |
url | Use a llama-server you run yourself; Cortana starts none |
Set HF_TOKEN in .env for gated repositories. The server log is .cortana/llama-server-<port>.log, and a server that dies is restarted on the next request.
Chat templates
llama-server renders the Jinja template inside the GGUF, and some are strict. Gemma 3's, for example, allows one leading system message and strictly alternating turns. llamacpp.chat_template.messages controls how Cortana's history is converted:
nativesends system, user, assistant and tool roles as is. Qwen, gpt-oss and Llama 3.x take it.alternatingmerges system messages and turns tool results into user turns, for Gemma 3 and older Mistral templates.auto(default) starts native and switches to alternating when the template rejects the turn order.
One model answers every role: helper, verifier and vision models named in the config use the loaded GGUF. Image generation stays on Ollama or Qwen-Image, and embeddings stay on Ollama unless llamacpp.embedding sets a model.
cortana.yml reference
The top-level sections. cortana.example.yml documents every key with comments.
| Section | Controls | Docs |
|---|---|---|
agent | Model, engine, reasoning, parallelism, tool iterations, permission mode, destructive-command confirmation, tool builder, context window, extra dirs, planning and verifier | Core, Planning |
memory | Provider, history length, working memory, semantic recall, chunking, Qdrant | Memory |
voice | TTS provider and mode, Kokoro/Qwen voices, HTTP server, hands-free listening | Voice |
image | Provider, size, steps, device, output folder, vision model | Images |
video | Wan model, size, frames, fps, steps | Videos |
heartbeat | Check-in interval and active hours | Heartbeat |
learning | Experience recipes | Learning |
ui | Show thinking, recaps, tips, helper model, /clear hint | CLI |
graphify | Knowledge graph, context mode, tools, excludes | Knowledge graph |
browser | Engine, channel, headless, idle timeout | Browser |
mcp | Enable, connect timeout, servers | MCP |
A typical starting point:
agent:
model: gpt-oss:20b
reasoning:
effort: medium
context: current_turn
max_parallel_tools: 6
max_tool_iterations: 10
permission_mode: auto # auto | plan | ask
confirm_destructive: true
context_window: 32768
planning: true
# extra_dirs: [~/Documents/notes]
memory:
provider: local # none | local | qdrant
semantic_recall:
enabled: true
embed_model: mxbai-embed-large:335m
voice:
provider: kokoro
ui:
show_thinking: true
recaps: true
tips: true
Environment variables
Every setting has an environment override, which wins over cortana.yml. The most useful:
| Variable | Purpose |
|---|---|
OLLAMA_MODEL | Chat model |
OLLAMA_HOST | Ollama server address |
CORTANA_ENGINE | ollama, docker-model-runner or llamacpp |
DOCKER_MODEL_RUNNER_URL | Docker Model Runner host URL |
OLLAMA_CONTEXT_WINDOW | Effective context size, for usage display and compaction |
OLLAMA_THINK | Turn model reasoning on or off |
OLLAMA_KEEP_ALIVE | How long the chat model stays loaded |
CONTEXT_COMPACTION_THRESHOLD | Compaction trigger, 0–1 (default 0.9) |
REASONING_EFFORT, REASONING_CONTEXT | Reasoning defaults |
PERMISSION_MODE | auto, plan or ask |
CONFIRM_DESTRUCTIVE | Ask before destructive commands |
CORTANA_CONFIG | Path to the YAML settings file |
CORTANA_EXTRA_DIRS | More folders the file tools may use |
CORTANA_STATE_DIR | Where schedules and background tasks are stored |
MEMORY_PROVIDER | none, local or qdrant |
MEMORY_EMBED_MODEL | Embedding model for recall |
QDRANT_URL, QDRANT_API_KEY | Qdrant server and credential |
TTS_PROVIDER, TTS_MODE | Voice provider (kokoro/qwen) and local/http |
KOKORO_TTS_VOICE | Kokoro voice |
IMAGE_PROVIDER, IMAGE_VISION_MODEL | Image provider and vision model |
BROWSER_ENABLED, BROWSER_HEADLESS | Browser tools |
HEARTBEAT_ENABLED | Self-started check-ins |
MCP_ENABLED | MCP servers |
A TTS_PROVIDER=qwen left in .env or your shell profile wins over voice.provider in cortana.yml. Remove it to get Kokoro back.
Running through Temporal
For a dashboard of workflow status and worker progress, run turns through Temporal. Start the server and UI with Docker, then a worker:
docker compose -f docker-compose.temporal.yml up -d
uv run --extra temporal python -m app.temporal_runner worker
Submit a normal turn from another terminal:
uv run --extra temporal python -m app.temporal_runner run \
--workspace "$PWD" "Explain how this project is structured"
Open http://localhost:8080 to see the durable workflow and activity history. The worker prints structured events (run_started, tool_started, run_finished, …). Set TEMPORAL_ADDRESS when the server runs on another machine.