Core Concepts
How an agent run works, how context stays small, and the services that make the agent aware of your code, proactive and better with use.
Core objects
Agentholds configuration and runs the model/tool loop.create_model_clientselects the configured engine: Ollama, Docker Model Runner, or llama.cpp. Each adapter maps requests to the server API and returns Cortana's shared response shape.Toolwraps a Python callable and its generated JSON schema.ToolGroupbundles tools the model loads on demand.Message,ToolCallandRunResultare a stable internal format, separate from provider response objects.AgentRunnerschedules complete agent runs under a semaphore and returns anAgentOutcomeper task.SubagentManageris model-driven delegation: it gives a parent agent tools to spawn, wait for, follow up and cancel subagents.MemoryProvideris a protocol for thread history and long-term context.Agent.memorytakes a provider, or a function that picks one per request.
Run lifecycle
user prompt + resource/thread identity
│
▼
resolve memory provider
├── load ordered short-term history
├── retrieve relevant long-term context
└── add the provider's tools (e.g. update_working_memory)
│
▼
system + history + per-turn context + user message
│
▼
model call
├── no tool calls ──► persist the new turn ──► RunResult
│
└── tool calls
├── authorize each call
├── validate arguments with Pydantic
├── run with bounded concurrency (default 4)
└── append results in the original request order
│
└──► next model call
The loop stops when the model replies without tool calls. run() and run_stream() share this lifecycle and return identically shaped history.
Past max_tool_iterations, the run raises MaxToolIterationsError rather than returning an empty answer. Pass on_max_iterations to ask first: returning True grants another max_tool_iterations.
async def ask(iterations: int) -> bool:
return input(f"Used {iterations} tool iterations. Continue? [y/N] ") == "y"
agent = Agent(name="dev", tools=[...], max_tool_iterations=10, on_max_iterations=ask)
Set parallel_tool_calls=False when calls in one response depend on each other, or tune max_parallel_tools.
Metrics
Every RunResult has RunMetrics: elapsed time, model and tool wall time, model calls, tool calls and errors, subagent spawns, and prompt/completion token counts. AgentLogger shows a spinner while the model thinks, a progress row per tool, and a compact summary line. Redirected output uses plain lines instead of animation.
Context compaction
When context reaches 90% of the window, the agent asks the same model for a continuation summary, replaces the older turns with it, and continues the current tool loop. The summary keeps the objective, constraints, decisions, paths, tool findings, test results, failures and remaining work. If the summary call fails, a bounded local fallback keeps the most recent transcript.
OLLAMA_CONTEXT_WINDOW=32768 CONTEXT_COMPACTION_THRESHOLD=0.85 cortana
Compaction needs the effective context size in OLLAMA_CONTEXT_WINDOW (or OLLAMA_CONTEXT_LENGTH, or agent.context_window). It stays off when the window is unknown. Use /compact [focus] to compact by hand.
Knowledge graph
KnowledgeGraph connects an agent to a graphify graph of the workspace's code: modules, classes, functions, imports and calls, grouped into communities. Building it needs no LLM.
uv tool install graphifyy
Without the CLI installed, the graph is simply left out. With it, the agent gets:
- Per-turn context. Symbols and files your prompt names (
Agent.run_stream,runtime.py) are explained withfile:linesources. Plain words never match, so "hi" costs nothing. - Structural tools.
query_knowledge_graph,explain_graph_node,graph_path,graph_affected("what breaks if I change this?"),graph_god_nodesandupdate_knowledge_graph. - Automatic rebuilds. After a turn that changed files, the graph rebuilds incrementally in the background.
from libs import Agent, KnowledgeGraph
graph = KnowledgeGraph("path/to/project", exclude=["vendor/"])
agent = Agent(name="dev", instructions="You are a coding assistant.", knowledge_graph=graph)
result = await agent.generate("What depends on SkillRegistry?")
await graph.aclose()
Configure it under graphify: in cortana.yml. context is lookup (default), report, both or off, and exclude is written to a managed block in .graphifyignore.
Heartbeat
The heartbeat lets the agent speak first. Every interval it runs one check-in turn with personality/HEARTBEAT.md as its checklist. A reply of just HEARTBEAT_OK is dropped; anything else is shown in the transcript, saved to the conversation, and spoken if voice mode is on.
heartbeat:
enabled: true
interval_minutes: 30
active_hours: "08:00-22:00" # local time; null for any hour
A beat is skipped while you're busy: a turn running, queued input, a half-typed prompt or an open approval dialog. Check-in tool calls still go through the permission gate. Use --no-heartbeat to turn it off for one session.
from libs import Heartbeat
heartbeat = Heartbeat(
check_in, # async (prompt) -> reply text
interval_seconds=30 * 60,
active_hours="08:00-22:00",
is_busy=lambda: ui.turn_running,
on_message=show_to_user,
)
await heartbeat.start()
Learning
Memory remembers what was said. Learning remembers what was done, so the agent gets to an answer faster the second time.
Experience
After a run with at least three tool calls, a background model call distills the trace into a short recipe: TASK, STEPS and AVOID. On a similar prompt later, the best recipes are shown to the model as a shortcut, with a reminder to check that the files it relies on haven't changed. Distilling happens after the reply, so it never slows one down.
/goodand/badrate the last answer. A prompt that starts like a correction ("no, …", "that's wrong") within 15 minutes also counts as bad.- Recipes rated bad twice with no good ratings stop being shown.
- A recipe rated good at least twice lets the run use less thinking (
trusted_think: low). /learnedlists recipes and shows whether runs with a recipe really use fewer tool calls.
from libs import Agent, Experience
experience = Experience(".cortana/experience.db", embedder="mxbai-embed-large:335m")
agent = Agent(name="assistant", tools=tools, experience=experience)
Teach mode
Do a task yourself while Cortana records your ! commands, notes and file changes, and it drafts a SKILL.md from them. See Shell & teach mode.
Repeated reads
Within a run, an identical repeat of a read-only call (read_file, grep_files, graph queries, …) isn't run again; the model is pointed at the earlier result. The cache clears after any write, command or MCP call.