Introduction
Cortana is a small async agent framework for local models, plus a terminal assistant built on it. Choose Ollama, Docker Model Runner, or llama.cpp as the model engine.
Overview
You can use Cortana in two ways:
- As an assistant. Run
cortanain any directory and chat with it in a full-screen terminal UI. It reads and edits files, runs commands, searches the web, drives a browser, and generates speech, images and video. Model inference can run locally through Ollama, Docker Model Runner, or llama.cpp. - As an SDK. Import
libsin Python to build your own agents. An agent combines a model, instructions, tools, and optional memory, voice and planning. The runtime owns the model/tool loop, andAgentRunnerruns many agents under a concurrency limit.
The code has two layers. src/libs/ is the provider-independent framework, and its public API is exported from libs. src/app/ is the CLI wiring: the terminal UI, slash commands, permissions and the runtime that owns every service.
Install
You need Python 3.12 (3.13 is not supported), uv, and one supported model engine. Ollama is the default; Docker Model Runner and llama.cpp are selectable in cortana.yml.
ollama pull gpt-oss:20b
To use Docker Model Runner instead, enable it in Docker Desktop, turn on host TCP access on port 12434, and pull a model:
docker model pull ai/smollm2
Then set agent.engine: docker-model-runner and agent.model: ai/smollm2. See Docker Model Runner setup for a complete configuration and embedding model.
Then launch from the repository:
scripts/cortana # syncs dependencies, then opens the TUI
scripts/cortana runs through uv run, so missing dependencies are installed before each launch, and Playwright's Chromium is installed on the first run. Plain uv sync && uv run python main.py also works for the default environment.
Start it from anywhere
To start Cortana by typing cortana, the way claude starts Claude Code, link the script onto your PATH once:
ln -s "$PWD/scripts/cortana" ~/.local/bin/cortana
cortana # full-screen TUI in the current directory
cortana --plain # line-by-line REPL
The directory you launch from becomes the workspace: the folder the file tools can read and write. If it has no cortana.yml of its own, the project's cortana.yml is used (unless CORTANA_CONFIG is set). .env is always read from the project.
Optional extras
| Feature | What to install |
|---|---|
| Voice (Kokoro TTS, Faster Whisper, microphone) | Installed by default. Kokoro also needs espeak-ng from your OS package manager. |
| Image generation (Qwen-Image-2.1) | Installed by default (the image group). |
| Video generation (Wan 2.1) | uv sync --extra video |
| Qwen TTS voices | cortana --qwen-tts creates and uses a separate .venv-qwen. |
Vision (view_image) | ollama pull qwen2.5vl:3b or configure a vision-capable model on your selected engine |
| Semantic memory recall | Ollama: ollama pull mxbai-embed-large:335m; DMR: docker model pull ai/nomic-embed-text-v2-moe |
| Code knowledge graph | uv tool install graphifyy |
Qwen TTS pins an older transformers than image and video generation need, which is why it lives in its own environment. See Voice & audio.
First run
On a new workspace, Cortana creates starter personality files and personality/BOOTSTRAP.md. On the first conversation it asks what to call itself, how it should sound, what to call you and how you like to work. When you confirm, it writes personality/USER.md and deletes the bootstrap file.
After that, just type. Some things to try:
› explain how this project is structured
› add a --verbose flag to cli.py and run the tests
› research the latest Ollama release and summarize what changed
› open news.ycombinator.com and tell me the top story
› draw a lighthouse at dusk, watercolor
› remind me every weekday at 9 to check the build
Press Shift+Tab to switch to plan mode before a larger change, so the agent investigates and writes a plan without editing anything. Esc interrupts a turn. Type /help for every command. Using the CLI covers all of this in detail.
Memory is configured in cortana.yml or with --memory local|qdrant. With memory on, conversations are saved and /resume brings them back.
Your first agent
The SDK is async. Create an agent, give it tools bound to a workspace, and run a prompt:
import asyncio
from pathlib import Path
from libs import Agent, create_builtin_tools
async def main() -> None:
agent = Agent(
name="coder",
model="gpt-oss:20b",
system_prompt=(
"You are a coding assistant. Inspect relevant files, make precise "
"edits, and run focused checks before answering."
),
tools=create_builtin_tools(Path.cwd()),
)
result = await agent.run("Explain how the package is structured")
print(result.output)
asyncio.run(main())
Agent.run() returns a RunResult with the final output, optional model thinking, the complete normalized message history, and RunMetrics.
Continue a conversation
History is explicit. Pass the previous result's messages into the next run:
first = await agent.run("Read pyproject.toml")
second = await agent.run("What Python version does it require?", history=first.messages)
The input list is copied, so your list is never mutated.
Give it memory
Or let a memory provider keep the history for you. resource scopes long-term memory (a user, a team, a project); thread is one conversation:
from libs import Agent, Memory
agent = Agent(name="assistant", model="gpt-oss:20b", memory=Memory(last_messages=20))
await agent.generate(
"Remember that releases happen on Fridays.",
memory={"thread": "planning", "resource": "project-acme"},
)
agent.stream(...) is the streaming version of generate. See Memory for working memory, semantic recall and Qdrant.
Custom tools
Decorate a typed sync or async function. The signature becomes the JSON schema, and Pydantic validates the model's arguments before the call:
from libs import tool
@tool
async def issue_status(issue_id: int, include_comments: bool = False) -> str:
"""Return the current status of an issue."""
...
Many agents at once
from libs import AgentRunner, AgentTask
outcomes = await AgentRunner(max_concurrency=2).run_many([
AgentTask(reviewer, "Review the API", task_id="review"),
AgentTask(tester, "Suggest test cases", task_id="tests"),
])
for outcome in outcomes:
print(outcome.task_id, outcome.result.output if outcome.succeeded else outcome.error)
Outcomes keep input order, and a failed task doesn't discard its siblings unless you pass fail_fast=True. The examples/ folder has runnable scripts for tools, voice, images and background agents.
Workspace & personality
The CLI shapes its behavior from Markdown files you can read and edit. AGENTS.md sits at the project root; personal context lives under personality/:
| File | Purpose | Loaded |
|---|---|---|
AGENTS.md | Operating and project instructions | Main agent and subagents |
personality/SOUL.md | Personality, boundaries and tone | Main agent |
personality/IDENTITY.md | Name and role | Main agent |
personality/USER.md | Your confirmed preferences | Main agent |
personality/MEMORY.md | Lessons saved with remember_lesson | Main agent, when present |
personality/HEARTBEAT.md | Checklist for self-started check-ins | Heartbeat turns only |
personality/BOOTSTRAP.md | One-time setup conversation | Until setup is confirmed |
The system prompt is rebuilt before every turn, so manual edits take effect immediately. Existing files are never overwritten; a missing core file in a partly configured workspace is marked in the prompt rather than recreated.
Self-learning. When you correct Cortana, it can call remember_lesson, which appends one dated sentence to personality/MEMORY.md. Lessons are capped at 400 characters, duplicates are skipped, and text that looks like a credential is rejected. Facts about you go to working memory instead; reusable procedures are learned as experience.
Keep personality files private if they hold personal context, and never put credentials in USER.md or MEMORY.md.
Safety
What the runtime enforces:
- Workspace scoping. File and search tools can't reach outside the workspace, including through
..or symlinks, unless you add a folder with--add-dir,/add-diroragent.extra_dirs. - Schema validation. Tool arguments are validated and unknown arguments are rejected.
- Atomic writes. File replacement is atomic, and writes to the same path are serialized.
- Bounded output. Reads, searches, listings and shell output are capped; shell timeouts kill the whole process group.
- Approvals. The CLI asks before destructive shell commands (
rm,sudo,git push,git reset --hard, …) and before creating or running a tool the model wrote, in every permission mode. - Least privilege.
create_read_only_tools()builds reviewer agents; subagent profiles can be narrowed but never broadened.
Trust boundary: run_command starts in the workspace but is not a sandbox. A shell command can reach anything the Python process can, including the network. For untrusted prompts or models, use ask mode or run the whole process in an OS or container sandbox.
For your own agents, authorize_tool is the hook for approvals and policy. See Tools.