C CORTANADocumentation › Guides

Core Concepts

How an agent run works, how context stays small, and the services that make the agent aware of your code, proactive and better with use.

Core objects

  • Agent holds configuration and runs the model/tool loop.
  • create_model_client selects the configured engine: Ollama, Docker Model Runner, or llama.cpp. Each adapter maps requests to the server API and returns Cortana's shared response shape.
  • Tool wraps a Python callable and its generated JSON schema. ToolGroup bundles tools the model loads on demand.
  • Message, ToolCall and RunResult are a stable internal format, separate from provider response objects.
  • AgentRunner schedules complete agent runs under a semaphore and returns an AgentOutcome per task.
  • SubagentManager is model-driven delegation: it gives a parent agent tools to spawn, wait for, follow up and cancel subagents.
  • MemoryProvider is a protocol for thread history and long-term context. Agent.memory takes a provider, or a function that picks one per request.

Run lifecycle

user prompt + resource/thread identity
    │
    ▼
resolve memory provider
    ├── load ordered short-term history
    ├── retrieve relevant long-term context
    └── add the provider's tools (e.g. update_working_memory)
    │
    ▼
system + history + per-turn context + user message
    │
    ▼
model call
    ├── no tool calls ──► persist the new turn ──► RunResult
    │
    └── tool calls
          ├── authorize each call
          ├── validate arguments with Pydantic
          ├── run with bounded concurrency (default 4)
          └── append results in the original request order
                       │
                       └──► next model call

The loop stops when the model replies without tool calls. run() and run_stream() share this lifecycle and return identically shaped history.

Past max_tool_iterations, the run raises MaxToolIterationsError rather than returning an empty answer. Pass on_max_iterations to ask first: returning True grants another max_tool_iterations.

async def ask(iterations: int) -> bool:
    return input(f"Used {iterations} tool iterations. Continue? [y/N] ") == "y"

agent = Agent(name="dev", tools=[...], max_tool_iterations=10, on_max_iterations=ask)

Set parallel_tool_calls=False when calls in one response depend on each other, or tune max_parallel_tools.

Metrics

Every RunResult has RunMetrics: elapsed time, model and tool wall time, model calls, tool calls and errors, subagent spawns, and prompt/completion token counts. AgentLogger shows a spinner while the model thinks, a progress row per tool, and a compact summary line. Redirected output uses plain lines instead of animation.

Context compaction

When context reaches 90% of the window, the agent asks the same model for a continuation summary, replaces the older turns with it, and continues the current tool loop. The summary keeps the objective, constraints, decisions, paths, tool findings, test results, failures and remaining work. If the summary call fails, a bounded local fallback keeps the most recent transcript.

OLLAMA_CONTEXT_WINDOW=32768 CONTEXT_COMPACTION_THRESHOLD=0.85 cortana

Compaction needs the effective context size in OLLAMA_CONTEXT_WINDOW (or OLLAMA_CONTEXT_LENGTH, or agent.context_window). It stays off when the window is unknown. Use /compact [focus] to compact by hand.

Knowledge graph

KnowledgeGraph connects an agent to a graphify graph of the workspace's code: modules, classes, functions, imports and calls, grouped into communities. Building it needs no LLM.

uv tool install graphifyy

Without the CLI installed, the graph is simply left out. With it, the agent gets:

  • Per-turn context. Symbols and files your prompt names (Agent.run_stream, runtime.py) are explained with file:line sources. Plain words never match, so "hi" costs nothing.
  • Structural tools. query_knowledge_graph, explain_graph_node, graph_path, graph_affected ("what breaks if I change this?"), graph_god_nodes and update_knowledge_graph.
  • Automatic rebuilds. After a turn that changed files, the graph rebuilds incrementally in the background.
from libs import Agent, KnowledgeGraph

graph = KnowledgeGraph("path/to/project", exclude=["vendor/"])
agent = Agent(name="dev", instructions="You are a coding assistant.", knowledge_graph=graph)
result = await agent.generate("What depends on SkillRegistry?")
await graph.aclose()

Configure it under graphify: in cortana.yml. context is lookup (default), report, both or off, and exclude is written to a managed block in .graphifyignore.

Heartbeat

The heartbeat lets the agent speak first. Every interval it runs one check-in turn with personality/HEARTBEAT.md as its checklist. A reply of just HEARTBEAT_OK is dropped; anything else is shown in the transcript, saved to the conversation, and spoken if voice mode is on.

heartbeat:
  enabled: true
  interval_minutes: 30
  active_hours: "08:00-22:00"   # local time; null for any hour

A beat is skipped while you're busy: a turn running, queued input, a half-typed prompt or an open approval dialog. Check-in tool calls still go through the permission gate. Use --no-heartbeat to turn it off for one session.

from libs import Heartbeat

heartbeat = Heartbeat(
    check_in,                        # async (prompt) -> reply text
    interval_seconds=30 * 60,
    active_hours="08:00-22:00",
    is_busy=lambda: ui.turn_running,
    on_message=show_to_user,
)
await heartbeat.start()

Learning

Memory remembers what was said. Learning remembers what was done, so the agent gets to an answer faster the second time.

Experience

After a run with at least three tool calls, a background model call distills the trace into a short recipe: TASK, STEPS and AVOID. On a similar prompt later, the best recipes are shown to the model as a shortcut, with a reminder to check that the files it relies on haven't changed. Distilling happens after the reply, so it never slows one down.

  • /good and /bad rate the last answer. A prompt that starts like a correction ("no, …", "that's wrong") within 15 minutes also counts as bad.
  • Recipes rated bad twice with no good ratings stop being shown.
  • A recipe rated good at least twice lets the run use less thinking (trusted_think: low).
  • /learned lists recipes and shows whether runs with a recipe really use fewer tool calls.
from libs import Agent, Experience

experience = Experience(".cortana/experience.db", embedder="mxbai-embed-large:335m")
agent = Agent(name="assistant", tools=tools, experience=experience)

Teach mode

Do a task yourself while Cortana records your ! commands, notes and file changes, and it drafts a SKILL.md from them. See Shell & teach mode.

Repeated reads

Within a run, an identical repeat of a read-only call (read_file, grep_files, graph queries, …) isn't run again; the model is pointed at the earlier result. The cache clears after any write, command or MCP call.