C CORTANADocumentation › Guides

Intelligence

Memory across conversations, delegation to subagents, plans the agent must actually finish, external tools over MCP, and work that runs in the background.

Memory

Memory follows Mastra's shape: one Memory object with pluggable storage, vector store and embedder. Pass a resource (who the memory belongs to) and a thread (one conversation) with each call.

from libs import Agent, Memory

agent = Agent(name="assistant", instructions="Be helpful.", memory=Memory())
await agent.generate(
    "Remember that releases happen on Fridays.",
    memory={"thread": "planning", "resource": "project-acme"},
)

The three memory types

  • Message history. Every message is saved, and the last last_messages (default 40) are replayed. The window moves in half-size steps so the configured model engine can reuse its prompt cache.
  • Working memory. A Markdown document shown on every turn, which the agent rewrites with update_working_memory. With scope="resource" (default) it follows you into every thread.
  • Semantic recall. Messages are embedded; each prompt pulls in the top_k most similar older messages, plus neighbours. If the embedding model is missing, recall is skipped with a warning.
from libs import LocalStore, LocalVector, Memory, MemoryConfig, SemanticRecall, WorkingMemory

memory = Memory(
    storage=LocalStore(".memory"),
    vector=LocalVector(".memory"),
    embedder="mxbai-embed-large:335m",
    options=MemoryConfig(
        last_messages=20,
        semantic_recall=SemanticRecall(top_k=3, message_range=2, scope="resource"),
        working_memory=WorkingMemory(enabled=True, scope="resource"),
    ),
)

Shortcuts work too: Memory(last_messages=20, semantic_recall=True). Other options: goals (add_goal/complete_goal tools), generate_title and read_only.

In the CLI

cortana --memory local --resource-id brian --thread-id cortana-development
memory:
  provider: local            # none | local | qdrant
  last_messages: 40
  working_scope: resource
  semantic_recall:
    enabled: true
    embed_model: mxbai-embed-large:335m
    top_k: 5

With Docker Model Runner, use a pulled embedding model such as ai/nomic-embed-text-v2-moe. It produces 768-dimensional vectors; use a separate collection such as cortana_memory_dmr or re-embed your existing collection when changing models.

/memory shows working memory, /resume lists saved conversations, /forget deletes the current one, and /reindex rebuilds recall vectors.

Qdrant

QdrantMemory keeps history, recall and ingested documents in one Qdrant collection. With Qdrant, relevant memories are added to each message automatically, so search_memory is loaded only when you refer to something not shown.

memory:
  provider: qdrant
  qdrant:
    url: http://192.168.3.230:6333          # https:// enables TLS
    collection: cortana_memory
    fallbacks: [http://localhost:6333, ":memory:"]
    probe_interval: 30

Keep QDRANT_API_KEY in .env. API keys are refused over plain http:// unless you opt in. With fallbacks, Cortana keeps working through a Qdrant outage and replays writes from .cortana/qdrant-outbox.json into the primary when it returns.

Documents are ingested with docling, chunked along their headings, and re-ingesting a source replaces its old chunks:

await memory.ingest_source("project-acme", "docs/architecture.pdf")
await memory.ingest_source("project-acme", "https://example.com/handbook")

Password-protected PDFs and Office files ask for their password (a masked prompt in the CLI). The password goes straight to decryption and never reaches the model, transcript or memory.

Custom providers

An agent only needs the MemoryProvider protocol, so SQL, Redis or any other store works:

class MemoryProvider(Protocol):
    async def build_context(self, resource_id, thread_id, prompt) -> list[Message]: ...
    async def load_history(self, resource_id, thread_id) -> list[Message]: ...
    async def persist(self, resource_id, thread_id, new_messages) -> None: ...

Subagents

Two ways to use more than one agent:

  • AgentRunner: your code decides which agents run, in parallel, with a batch-wide limit.
  • SubagentManager: the model decides. It picks a host-approved profile and hands it a task.
from libs import Agent, SubagentDefinition, SubagentManager, create_builtin_tools, create_read_only_tools

manager = SubagentManager(
    parent_model="gpt-oss:20b",
    max_concurrency=3,
    definitions=[
        SubagentDefinition(name="explorer", description="Read-only repository research",
                           system_prompt="Investigate independently and cite files.",
                           tools=create_read_only_tools()),
        SubagentDefinition(name="worker", description="Implementation and verification",
                           system_prompt="Implement the task and run focused checks.",
                           tools=create_builtin_tools()),
    ],
)
parent = Agent(
    name="lead",
    model="gpt-oss:20b",
    system_prompt="\n\n".join(["Lead the task.", manager.prompt_instructions]),
    tools=[*create_builtin_tools(), manager.delegate_tool, *manager.tools],
)

delegate(agent_name, task) starts a subagent, waits and returns its result; small local models use it far more reliably than spawn-and-wait. For parallel work the parent also gets spawn_subagent, wait_for_subagents, subagent_status, follow_up_subagent and cancel_subagent. A parent can narrow a profile's tools but never add to them. Call await manager.close() when the session ends.

The CLI ships an explorer (read-only, with web search) and a worker (coding tools). The agent uses the explorer before changing code it hasn't read, keeping its own context small. In plan mode the explorer is allowed and the worker is refused. ← on an empty prompt and /agents show running subagents.

Planning & verification

Small local models often describe a multi-step job, do part of it and report it finished. Planning gives each run a checklist it must work through, and verification checks that checklist against the tool calls that actually ran.

from libs import Agent, TaskVerifier

agent = Agent(
    name="dev",
    model="gpt-oss:20b",
    tools=[...],
    planning=True,                              # adds update_plan
    plan_dir=".cortana/plans",                  # keep plans as Markdown
    verifier=TaskVerifier("gpt-oss:20b", rounds=2),
    max_tool_iterations=25,                     # plans take extra calls
)

The model writes a plan (goal, findings, steps with done_when checks, verification commands), then marks steps off:

Plan (1/3 done):
[x] 1. Create calc.py
[~] 2. Create test_calc.py
[ ] 3. Run the tests
Now do step 2: Create test_calc.py
  • Plan before changing files. The first write_file, edit_file or create_tool call of a run with no plan is sent back, asking for a plan that covers every requirement.
  • No unfinished steps. A reply while steps are still open is sent back with the list.
  • Plans outlive the run. With a plan_dir, a thread's unfinished plan is picked up again when you say "go ahead". You can edit the Markdown file between runs.

The verifier catches "said, not done": steps marked done with no tool call since, files claimed but never written, failures reported as success, commands written out instead of run, and "it works" with no check after the last edit. An optional small inspector model opens the files a turn made and checks each requirement.

agent:
  planning: true
  verifier:
    model: gpt-oss:20b      # null: model-free checks only
    rounds: 2
    inspector:
      model: granite4:7b-a1b-h

In the CLI, use plan mode to review a plan before it runs; see Permission modes.

MCP servers

Cortana connects to Model Context Protocol servers over stdio, Streamable HTTP or legacy SSE. Each tool is named <server>_<tool> and goes through the same validation, permissions and logging as built-in tools. By default each server's tools form one on-demand group, keeping prompts small for local models.

Servers come from three places; later ones win on a name clash:

  1. ~/.cortana/mcp.json: user scope, every workspace.
  2. <workspace>/.mcp.json: project scope, the same format as Claude Code's.
  3. mcp.servers in cortana.yml.
{
  "mcpServers": {
    "playwright": {"type": "stdio", "command": "npx", "args": ["-y", "@playwright/mcp@latest"]},
    "postman": {
      "type": "http",
      "url": "https://mcp.postman.com/minimal",
      "headers": {"Authorization": "Bearer ${POSTMAN_API_KEY}"}
    }
  }
}

${VAR} and ${VAR:-default} are expanded from .env or the environment; a server whose variable is unset is skipped with a reason in /mcp.

KeyDefaultMeaning
enabledtruefalse keeps the entry but doesn't connect
deferredtruefalse sends the tools every turn
tools / exclude_toolsallAllow- or deny-list of tool names
prefixserver nameTool-name prefix; "" disables it
call_timeout120Seconds per tool call
max_output_chars20000Longer results are truncated

Tools the server marks read-only always run. In ask mode every other MCP tool needs approval; plan mode refuses them. MCP servers run with your privileges, so only add ones you trust. stdio logs go to .cortana/mcp/<name>.log. Add servers from the command line with cortana mcp add.

Background tasks

BackgroundTaskManager runs opted-in tool calls outside the agent loop, with state in SQLite. It handles concurrency limits, timeouts, retries, cancellation and recovery after a restart.

from libs import Agent, BackgroundTaskManager, background_tool

@background_tool(timeout_seconds=600, max_retries=1)
async def research(topic: str) -> str:
    return await run_long_research(topic)

manager = BackgroundTaskManager(".cortana/background_tasks.db", on_finished=report)
agent = Agent(name="researcher", tools=[research], background_task_manager=manager)

A deferred tool returns a task ID immediately and the turn continues; awaited tools are tracked while the turn waits. Only the developer can make a tool eligible. on_finished is called exactly once per deferred result, even across restarts.

In the CLI, video generation runs this way. The status line shows ◐ 1 background task · generate_video 2m 14s, /tasks lists tasks, /tasks cancel <id> stops one, and the agent tells you when a task finishes.

Artifacts

ArtifactStore keeps what the agent extracted or made outside its context window, in <workspace>/.cortana/artifacts.db.

  • Documents. fetch_url saves each document's full text and shows the model an 8,000-character preview with the artifact id. An unchanged local file is never extracted twice.
  • Made files. Speech, images, videos, screenshots and pages are recorded with what they were made from, so "play the life plan audio again" becomes a lookup instead of a folder search.

The model reads the store with list_artifacts(query, kind) and read_artifact(id, start, find). Delete artifacts.db to clear it.

from libs import Agent, ArtifactStore, create_builtin_tools

artifacts = ArtifactStore(".cortana/artifacts.db")
agent = Agent(name="assistant", tools=create_builtin_tools(".", artifacts=artifacts))