C CORTANADocumentation › Guides

Introduction

Cortana is a small async agent framework for local models, plus a terminal assistant built on it. Choose Ollama, Docker Model Runner, or llama.cpp as the model engine.

Overview

You can use Cortana in two ways:

  • As an assistant. Run cortana in any directory and chat with it in a full-screen terminal UI. It reads and edits files, runs commands, searches the web, drives a browser, and generates speech, images and video. Model inference can run locally through Ollama, Docker Model Runner, or llama.cpp.
  • As an SDK. Import libs in Python to build your own agents. An agent combines a model, instructions, tools, and optional memory, voice and planning. The runtime owns the model/tool loop, and AgentRunner runs many agents under a concurrency limit.

The code has two layers. src/libs/ is the provider-independent framework, and its public API is exported from libs. src/app/ is the CLI wiring: the terminal UI, slash commands, permissions and the runtime that owns every service.

Install

You need Python 3.12 (3.13 is not supported), uv, and one supported model engine. Ollama is the default; Docker Model Runner and llama.cpp are selectable in cortana.yml.

ollama pull gpt-oss:20b

To use Docker Model Runner instead, enable it in Docker Desktop, turn on host TCP access on port 12434, and pull a model:

docker model pull ai/smollm2

Then set agent.engine: docker-model-runner and agent.model: ai/smollm2. See Docker Model Runner setup for a complete configuration and embedding model.

Then launch from the repository:

scripts/cortana            # syncs dependencies, then opens the TUI

scripts/cortana runs through uv run, so missing dependencies are installed before each launch, and Playwright's Chromium is installed on the first run. Plain uv sync && uv run python main.py also works for the default environment.

Start it from anywhere

To start Cortana by typing cortana, the way claude starts Claude Code, link the script onto your PATH once:

ln -s "$PWD/scripts/cortana" ~/.local/bin/cortana
cortana                    # full-screen TUI in the current directory
cortana --plain            # line-by-line REPL

The directory you launch from becomes the workspace: the folder the file tools can read and write. If it has no cortana.yml of its own, the project's cortana.yml is used (unless CORTANA_CONFIG is set). .env is always read from the project.

Optional extras

FeatureWhat to install
Voice (Kokoro TTS, Faster Whisper, microphone)Installed by default. Kokoro also needs espeak-ng from your OS package manager.
Image generation (Qwen-Image-2.1)Installed by default (the image group).
Video generation (Wan 2.1)uv sync --extra video
Qwen TTS voicescortana --qwen-tts creates and uses a separate .venv-qwen.
Vision (view_image)ollama pull qwen2.5vl:3b or configure a vision-capable model on your selected engine
Semantic memory recallOllama: ollama pull mxbai-embed-large:335m; DMR: docker model pull ai/nomic-embed-text-v2-moe
Code knowledge graphuv tool install graphifyy

Qwen TTS pins an older transformers than image and video generation need, which is why it lives in its own environment. See Voice & audio.

First run

On a new workspace, Cortana creates starter personality files and personality/BOOTSTRAP.md. On the first conversation it asks what to call itself, how it should sound, what to call you and how you like to work. When you confirm, it writes personality/USER.md and deletes the bootstrap file.

After that, just type. Some things to try:

› explain how this project is structured
› add a --verbose flag to cli.py and run the tests
› research the latest Ollama release and summarize what changed
› open news.ycombinator.com and tell me the top story
› draw a lighthouse at dusk, watercolor
› remind me every weekday at 9 to check the build

Press Shift+Tab to switch to plan mode before a larger change, so the agent investigates and writes a plan without editing anything. Esc interrupts a turn. Type /help for every command. Using the CLI covers all of this in detail.

Memory is configured in cortana.yml or with --memory local|qdrant. With memory on, conversations are saved and /resume brings them back.

Your first agent

The SDK is async. Create an agent, give it tools bound to a workspace, and run a prompt:

import asyncio
from pathlib import Path

from libs import Agent, create_builtin_tools


async def main() -> None:
    agent = Agent(
        name="coder",
        model="gpt-oss:20b",
        system_prompt=(
            "You are a coding assistant. Inspect relevant files, make precise "
            "edits, and run focused checks before answering."
        ),
        tools=create_builtin_tools(Path.cwd()),
    )
    result = await agent.run("Explain how the package is structured")
    print(result.output)


asyncio.run(main())

Agent.run() returns a RunResult with the final output, optional model thinking, the complete normalized message history, and RunMetrics.

Continue a conversation

History is explicit. Pass the previous result's messages into the next run:

first = await agent.run("Read pyproject.toml")
second = await agent.run("What Python version does it require?", history=first.messages)

The input list is copied, so your list is never mutated.

Give it memory

Or let a memory provider keep the history for you. resource scopes long-term memory (a user, a team, a project); thread is one conversation:

from libs import Agent, Memory

agent = Agent(name="assistant", model="gpt-oss:20b", memory=Memory(last_messages=20))
await agent.generate(
    "Remember that releases happen on Fridays.",
    memory={"thread": "planning", "resource": "project-acme"},
)

agent.stream(...) is the streaming version of generate. See Memory for working memory, semantic recall and Qdrant.

Custom tools

Decorate a typed sync or async function. The signature becomes the JSON schema, and Pydantic validates the model's arguments before the call:

from libs import tool


@tool
async def issue_status(issue_id: int, include_comments: bool = False) -> str:
    """Return the current status of an issue."""
    ...

Many agents at once

from libs import AgentRunner, AgentTask

outcomes = await AgentRunner(max_concurrency=2).run_many([
    AgentTask(reviewer, "Review the API", task_id="review"),
    AgentTask(tester, "Suggest test cases", task_id="tests"),
])
for outcome in outcomes:
    print(outcome.task_id, outcome.result.output if outcome.succeeded else outcome.error)

Outcomes keep input order, and a failed task doesn't discard its siblings unless you pass fail_fast=True. The examples/ folder has runnable scripts for tools, voice, images and background agents.

Workspace & personality

The CLI shapes its behavior from Markdown files you can read and edit. AGENTS.md sits at the project root; personal context lives under personality/:

FilePurposeLoaded
AGENTS.mdOperating and project instructionsMain agent and subagents
personality/SOUL.mdPersonality, boundaries and toneMain agent
personality/IDENTITY.mdName and roleMain agent
personality/USER.mdYour confirmed preferencesMain agent
personality/MEMORY.mdLessons saved with remember_lessonMain agent, when present
personality/HEARTBEAT.mdChecklist for self-started check-insHeartbeat turns only
personality/BOOTSTRAP.mdOne-time setup conversationUntil setup is confirmed

The system prompt is rebuilt before every turn, so manual edits take effect immediately. Existing files are never overwritten; a missing core file in a partly configured workspace is marked in the prompt rather than recreated.

Self-learning. When you correct Cortana, it can call remember_lesson, which appends one dated sentence to personality/MEMORY.md. Lessons are capped at 400 characters, duplicates are skipped, and text that looks like a credential is rejected. Facts about you go to working memory instead; reusable procedures are learned as experience.

Keep personality files private if they hold personal context, and never put credentials in USER.md or MEMORY.md.

Safety

What the runtime enforces:

  • Workspace scoping. File and search tools can't reach outside the workspace, including through .. or symlinks, unless you add a folder with --add-dir, /add-dir or agent.extra_dirs.
  • Schema validation. Tool arguments are validated and unknown arguments are rejected.
  • Atomic writes. File replacement is atomic, and writes to the same path are serialized.
  • Bounded output. Reads, searches, listings and shell output are capped; shell timeouts kill the whole process group.
  • Approvals. The CLI asks before destructive shell commands (rm, sudo, git push, git reset --hard, …) and before creating or running a tool the model wrote, in every permission mode.
  • Least privilege. create_read_only_tools() builds reviewer agents; subagent profiles can be narrowed but never broadened.

Trust boundary: run_command starts in the workspace but is not a sandbox. A shell command can reach anything the Python process can, including the network. For untrusted prompts or models, use ask mode or run the whole process in an OS or container sandbox.

For your own agents, authorize_tool is the hook for approvals and policy. See Tools.