Intelligence
Memory across conversations, delegation to subagents, plans the agent must actually finish, external tools over MCP, and work that runs in the background.
Memory
Memory follows Mastra's shape: one Memory object with pluggable storage, vector store and embedder. Pass a resource (who the memory belongs to) and a thread (one conversation) with each call.
from libs import Agent, Memory
agent = Agent(name="assistant", instructions="Be helpful.", memory=Memory())
await agent.generate(
"Remember that releases happen on Fridays.",
memory={"thread": "planning", "resource": "project-acme"},
)
The three memory types
- Message history. Every message is saved, and the last
last_messages(default 40) are replayed. The window moves in half-size steps so the configured model engine can reuse its prompt cache. - Working memory. A Markdown document shown on every turn, which the agent rewrites with
update_working_memory. Withscope="resource"(default) it follows you into every thread. - Semantic recall. Messages are embedded; each prompt pulls in the
top_kmost similar older messages, plus neighbours. If the embedding model is missing, recall is skipped with a warning.
from libs import LocalStore, LocalVector, Memory, MemoryConfig, SemanticRecall, WorkingMemory
memory = Memory(
storage=LocalStore(".memory"),
vector=LocalVector(".memory"),
embedder="mxbai-embed-large:335m",
options=MemoryConfig(
last_messages=20,
semantic_recall=SemanticRecall(top_k=3, message_range=2, scope="resource"),
working_memory=WorkingMemory(enabled=True, scope="resource"),
),
)
Shortcuts work too: Memory(last_messages=20, semantic_recall=True). Other options: goals (add_goal/complete_goal tools), generate_title and read_only.
In the CLI
cortana --memory local --resource-id brian --thread-id cortana-development
memory:
provider: local # none | local | qdrant
last_messages: 40
working_scope: resource
semantic_recall:
enabled: true
embed_model: mxbai-embed-large:335m
top_k: 5
With Docker Model Runner, use a pulled embedding model such as ai/nomic-embed-text-v2-moe. It produces 768-dimensional vectors; use a separate collection such as cortana_memory_dmr or re-embed your existing collection when changing models.
/memory shows working memory, /resume lists saved conversations, /forget deletes the current one, and /reindex rebuilds recall vectors.
Qdrant
QdrantMemory keeps history, recall and ingested documents in one Qdrant collection. With Qdrant, relevant memories are added to each message automatically, so search_memory is loaded only when you refer to something not shown.
memory:
provider: qdrant
qdrant:
url: http://192.168.3.230:6333 # https:// enables TLS
collection: cortana_memory
fallbacks: [http://localhost:6333, ":memory:"]
probe_interval: 30
Keep QDRANT_API_KEY in .env. API keys are refused over plain http:// unless you opt in. With fallbacks, Cortana keeps working through a Qdrant outage and replays writes from .cortana/qdrant-outbox.json into the primary when it returns.
Documents are ingested with docling, chunked along their headings, and re-ingesting a source replaces its old chunks:
await memory.ingest_source("project-acme", "docs/architecture.pdf")
await memory.ingest_source("project-acme", "https://example.com/handbook")
Password-protected PDFs and Office files ask for their password (a masked prompt in the CLI). The password goes straight to decryption and never reaches the model, transcript or memory.
Custom providers
An agent only needs the MemoryProvider protocol, so SQL, Redis or any other store works:
class MemoryProvider(Protocol):
async def build_context(self, resource_id, thread_id, prompt) -> list[Message]: ...
async def load_history(self, resource_id, thread_id) -> list[Message]: ...
async def persist(self, resource_id, thread_id, new_messages) -> None: ...
Subagents
Two ways to use more than one agent:
AgentRunner: your code decides which agents run, in parallel, with a batch-wide limit.SubagentManager: the model decides. It picks a host-approved profile and hands it a task.
from libs import Agent, SubagentDefinition, SubagentManager, create_builtin_tools, create_read_only_tools
manager = SubagentManager(
parent_model="gpt-oss:20b",
max_concurrency=3,
definitions=[
SubagentDefinition(name="explorer", description="Read-only repository research",
system_prompt="Investigate independently and cite files.",
tools=create_read_only_tools()),
SubagentDefinition(name="worker", description="Implementation and verification",
system_prompt="Implement the task and run focused checks.",
tools=create_builtin_tools()),
],
)
parent = Agent(
name="lead",
model="gpt-oss:20b",
system_prompt="\n\n".join(["Lead the task.", manager.prompt_instructions]),
tools=[*create_builtin_tools(), manager.delegate_tool, *manager.tools],
)
delegate(agent_name, task) starts a subagent, waits and returns its result; small local models use it far more reliably than spawn-and-wait. For parallel work the parent also gets spawn_subagent, wait_for_subagents, subagent_status, follow_up_subagent and cancel_subagent. A parent can narrow a profile's tools but never add to them. Call await manager.close() when the session ends.
The CLI ships an explorer (read-only, with web search) and a worker (coding tools). The agent uses the explorer before changing code it hasn't read, keeping its own context small. In plan mode the explorer is allowed and the worker is refused. ← on an empty prompt and /agents show running subagents.
Planning & verification
Small local models often describe a multi-step job, do part of it and report it finished. Planning gives each run a checklist it must work through, and verification checks that checklist against the tool calls that actually ran.
from libs import Agent, TaskVerifier
agent = Agent(
name="dev",
model="gpt-oss:20b",
tools=[...],
planning=True, # adds update_plan
plan_dir=".cortana/plans", # keep plans as Markdown
verifier=TaskVerifier("gpt-oss:20b", rounds=2),
max_tool_iterations=25, # plans take extra calls
)
The model writes a plan (goal, findings, steps with done_when checks, verification commands), then marks steps off:
Plan (1/3 done):
[x] 1. Create calc.py
[~] 2. Create test_calc.py
[ ] 3. Run the tests
Now do step 2: Create test_calc.py
- Plan before changing files. The first
write_file,edit_fileorcreate_toolcall of a run with no plan is sent back, asking for a plan that covers every requirement. - No unfinished steps. A reply while steps are still open is sent back with the list.
- Plans outlive the run. With a
plan_dir, a thread's unfinished plan is picked up again when you say "go ahead". You can edit the Markdown file between runs.
The verifier catches "said, not done": steps marked done with no tool call since, files claimed but never written, failures reported as success, commands written out instead of run, and "it works" with no check after the last edit. An optional small inspector model opens the files a turn made and checks each requirement.
agent:
planning: true
verifier:
model: gpt-oss:20b # null: model-free checks only
rounds: 2
inspector:
model: granite4:7b-a1b-h
In the CLI, use plan mode to review a plan before it runs; see Permission modes.
MCP servers
Cortana connects to Model Context Protocol servers over stdio, Streamable HTTP or legacy SSE. Each tool is named <server>_<tool> and goes through the same validation, permissions and logging as built-in tools. By default each server's tools form one on-demand group, keeping prompts small for local models.
Servers come from three places; later ones win on a name clash:
~/.cortana/mcp.json: user scope, every workspace.<workspace>/.mcp.json: project scope, the same format as Claude Code's.mcp.serversincortana.yml.
{
"mcpServers": {
"playwright": {"type": "stdio", "command": "npx", "args": ["-y", "@playwright/mcp@latest"]},
"postman": {
"type": "http",
"url": "https://mcp.postman.com/minimal",
"headers": {"Authorization": "Bearer ${POSTMAN_API_KEY}"}
}
}
}
${VAR} and ${VAR:-default} are expanded from .env or the environment; a server whose variable is unset is skipped with a reason in /mcp.
| Key | Default | Meaning |
|---|---|---|
enabled | true | false keeps the entry but doesn't connect |
deferred | true | false sends the tools every turn |
tools / exclude_tools | all | Allow- or deny-list of tool names |
prefix | server name | Tool-name prefix; "" disables it |
call_timeout | 120 | Seconds per tool call |
max_output_chars | 20000 | Longer results are truncated |
Tools the server marks read-only always run. In ask mode every other MCP tool needs approval; plan mode refuses them. MCP servers run with your privileges, so only add ones you trust. stdio logs go to .cortana/mcp/<name>.log. Add servers from the command line with cortana mcp add.
Background tasks
BackgroundTaskManager runs opted-in tool calls outside the agent loop, with state in SQLite. It handles concurrency limits, timeouts, retries, cancellation and recovery after a restart.
from libs import Agent, BackgroundTaskManager, background_tool
@background_tool(timeout_seconds=600, max_retries=1)
async def research(topic: str) -> str:
return await run_long_research(topic)
manager = BackgroundTaskManager(".cortana/background_tasks.db", on_finished=report)
agent = Agent(name="researcher", tools=[research], background_task_manager=manager)
A deferred tool returns a task ID immediately and the turn continues; awaited tools are tracked while the turn waits. Only the developer can make a tool eligible. on_finished is called exactly once per deferred result, even across restarts.
In the CLI, video generation runs this way. The status line shows ◐ 1 background task · generate_video 2m 14s, /tasks lists tasks, /tasks cancel <id> stops one, and the agent tells you when a task finishes.
Artifacts
ArtifactStore keeps what the agent extracted or made outside its context window, in <workspace>/.cortana/artifacts.db.
- Documents.
fetch_urlsaves each document's full text and shows the model an 8,000-character preview with the artifact id. An unchanged local file is never extracted twice. - Made files. Speech, images, videos, screenshots and pages are recorded with what they were made from, so "play the life plan audio again" becomes a lookup instead of a folder search.
The model reads the store with list_artifacts(query, kind) and read_artifact(id, start, find). Delete artifacts.db to clear it.
from libs import Agent, ArtifactStore, create_builtin_tools
artifacts = ArtifactStore(".cortana/artifacts.db")
agent = Agent(name="assistant", tools=create_builtin_tools(".", artifacts=artifacts))