Capabilities
What the agent can do: built-in tools, a real browser, skills, voice, image and video generation, and schedules.
Tools
Tool factories bind every tool to an explicit workspace:
from libs import create_builtin_tools, create_read_only_tools
coding_tools = create_builtin_tools("/work/project")
review_tools = create_read_only_tools("/work/project")
| Kind | Tools |
|---|---|
| Environment | get_system_info, get_datetime, get_device_context |
| Files | read_file, write_file, edit_file |
| Search | glob_files, grep_files, list_directory |
| Shell | run_command |
| Web | search, research, fetch_url, show_media, image/video search, downloads |
The read-only set leaves out file writes and the shell but keeps web research, which suits explorer and reviewer agents.
searchcollects up to 50 results across engines and opens with a breakdown by kind, site and recurring theme, so the model explains sources instead of pasting links.research(query, pages=5)searches, then reads the top pages in parallel and returns numbered sources to cite as[n].fetch_urlturns pages and documents (PDF, DOCX, PPTX, …) into clean Markdown with code snippets, images and media, and transcribes audio and video. Pages that come back empty are rendered in the browser.
Loading tools on demand
Every tool schema is sent with every request, so put rarely needed tools in a ToolGroup. The agent gets a load_tools tool listing each group; a group joins the run when the model loads it or calls one of its tools by name.
from libs import Agent, ToolGroup, create_audio_tools
agent = Agent(
name="assistant",
tools=core_tools,
deferred_tools=[ToolGroup("audio", "speak aloud and save speech audio", create_audio_tools(root))],
)
The CLI keeps files, shell, web search, workspace, knowledge-graph and skill tools always on, and defers audio, images, media, subagents, schedules, git, tool_builder and each MCP server. /tools lists them.
Authorization
authorize_tool runs before every call. Return True to allow, or False or a reason to deny; the model gets the reason and can recover.
def review_policy(call: ToolCall) -> bool | str:
if call.name in {"write_file", "edit_file", "run_command"}:
return "review agents are read-only"
return True
agent = Agent(name="reviewer", tools=create_builtin_tools(), authorize_tool=review_policy)
Git
The deferred git group gives the agent git_status, git_diff, git_log, git_show, git_branch and git_commit for the workspace repository. They run git with an argument list, not a shell, and refuse refs and paths that start with -.
git_commit stages exactly the paths it is given and commits only them, leaving other staged or unstaged work alone. Push, reset and rebase are not included. plan and ask modes treat git_branch and git_commit as edits. In the SDK, create_git_tools(workspace) returns the same tools.
Tools the agent writes itself
When no tool fits, the agent can write one with create_tool(name, code, description, dependencies). It runs in a separate, time-limited Python subprocess with its own site-packages, so your environment is never changed. save_created_tool copies it into the workspace as a standalone script you can run with uv run. The CLI always asks before creating or running one.
The subprocess is not a sandbox: a created tool can do anything the Python process can. Review the code in the approval prompt.
Browser
The agent drives a real browser through Python Playwright. It reads each page as an accessibility snapshot whose elements carry refs, and acts on those refs instead of guessing selectors:
browser_open("news.ycombinator.com")
Page: Hacker News
Snapshot (act on elements by ref, e.g. browser_click(ref='e12')):
- link "new" [ref=e17] -> newest
- link "147 comments" [ref=e51] -> item?id=49823582
...
browser_click(ref="e51")
Tools: browser_open (always on), browser_snapshot, browser_click, browser_type, browser_select, browser_press, browser_back, browser_wait, browser_read, browser_screenshot and browser_close. Opening a workspace .html file serves it locally so its CSS and scripts load, which is handy for checking pages the agent builds.
Compared with the Playwright MCP server, the built-in browser sends far smaller schemas and snapshots and has no arbitrary-code tool, which matters for small local models. Clicks, typing, selection and key presses are refused in plan mode and ask first in ask mode; the agent is also told to ask before buying, posting, sending or deleting anything.
browser:
enabled: true
engine: chromium # chromium | firefox | webkit
channel: "" # or chrome, msedge to use an installed browser
headless: true # false shows the window
idle_timeout: 900 # seconds before an idle browser closes
uv run playwright install chromium # once (scripts/cortana does this for you)
Skills
A skill is a folder with a SKILL.md: frontmatter plus task instructions. Skills live under skills/ or .agents/skills/.
---
name: release-check
description: Verify a release candidate before publishing.
---
Read the changelog, run `{baseDir}/check.sh`, and summarize failures.
Only the frontmatter is read at startup. The model finds skills with find_skills(query) and reads one with load_skill(name). Each turn, up to two clearly matching skills are suggested. A skill written for Cortana (agent: cortana) is included outright when it ranks highly or when the request contains one of its triggers:
triggers: [website, web page, landing page, html, css]
agent: cortana
Skills written for other agents often name tools like Bash or WebFetch; load_skill adds a note mapping each to its equivalent here. /skills [query] searches them in the CLI, and teach mode creates new ones.
Treat skill instructions as code: inspect third-party skills before installing them.
Voice & audio
Speech output uses Kokoro-82M (default) or Qwen3-TTS; speech input uses Faster Whisper with voice activity detection. Everything runs locally, or TTS can stream from your own HTTP server.
Talk to Cortana
- Press Ctrl+R or type
/voice, or start withcortana --listen. Pause to send; replies are spoken sentence by sentence as the model writes them. - Esc stops a spoken reply; say "stop listening" to turn it off.
- On macOS, give your terminal microphone access under System Settings → Privacy & Security → Microphone.
- Ask "use kokoro" or "use qwen" to switch voices mid-session, if that provider is installed.
| To | Do |
|---|---|
| Start with the configured provider | cortana |
| Force Kokoro | cortana --kokoro-tts |
Force Qwen (separate .venv-qwen) | cortana --qwen-tts |
| Speak through your TTS server | cortana --tts http |
voice:
provider: kokoro # qwen | kokoro
mode: local # local | http
kokoro:
voice: af_heart
language: a # a = American English, b = British English, ...
http:
url: http://my-server.local:8880/v1/audio/speech
api_style: openai
fallback_after: 3 # seconds; speak locally if the server is slow
listen:
barge_in: false # true: talk over replies to interrupt
Emojis and tags like [laughs] or [whispers] are acted out rather than read aloud. Qwen can clone a voice from 5–15 seconds of clean speech; only clone voices you have permission to use.
In your own code
from libs import Agent, SynthesizerVoice, VoiceConversation, play_audio
agent = Agent(name="assistant", model="gpt-oss:20b", voice=SynthesizerVoice())
reply = await agent.run("Say hello in one short sentence.")
await play_audio(await agent.speak(reply.output))
await VoiceConversation(agent).run() # hands-free until "stop listening"
Voice providers follow Mastra's shape (agent.speak, agent.listen, get_speakers), and VoicePipeline chains speech-to-text, the agent and streamed speech. docs/audio.md covers custom providers, cloning, realtime speech and every setting.
Images
Cortana generates and edits images, and can look at them with a vision-capable model to check its own edits. The vision model is served by the configured chat engine.
| Provider | Model | Notes |
|---|---|---|
qwen (default) | Qwen-Image-2.1 via diffusers | Local CUDA, Apple MPS or CPU; renders text well; supports editing |
ollama | x/z-image-turbo, x/flux2-klein, … | Runs on your Ollama server |
- Ask in plain language ("draw a lighthouse at dusk, watercolor", then "make the sky orange"), or use
/image <prompt>and/image edit <change>. edit_imagekeeps the original, then asks the vision model whether the change is really there, so the agent can try again.view_imagedescribes an image or answers a question about it. With Ollama, the default isqwen2.5vl:3b(ollama pull qwen2.5vl:3b). With Docker Model Runner, select a DMR model that supports image input; the examplesmollm2and Nomic embedding models are text-only.
Images are saved as <timestamp>-<prompt>-AI-generated.png in ~/Pictures unless you set image.output_dir. The first generation downloads the model into model/image/. On a 32 GB Apple silicon Mac, 1024×1024 at 40 steps takes about 4.5 minutes; 512×512 at 12 steps takes under one.
from libs import ImageConfig, ImageGenerator
images = ImageGenerator(ImageConfig(provider="qwen"), workspace=".")
result = await images.generate('A neon shop sign that reads "CORTANA", rainy night', seed=42)
edited = await images.generate("Change the background to a sunset beach", images=[result.path])
Videos
Short clips come from Wan 2.1 T2V 1.3B running locally through diffusers. Install it with uv sync --extra video; the first clip downloads about 17.6 GB into model/video/.
Ask for a clip ("a fox running through snow, low tracking shot"). generate_video runs as a deferred background task, so you can keep working; the agent tells you when the MP4 is ready. The default is 832×480, 33 frames at 16 fps (about two seconds); raise num_frames to 81 for about five.
video:
output_dir: videos
width: 832
height: 480
num_frames: 33 # 4k+1; 81 is the model's maximum
fps: 16
steps: 30
Schedules
Ask Cortana to do something on a schedule ("every weekday at 9, summarize my open PRs"), and it creates a durable cron schedule. Schedules live in SQLite, survive restarts, and their results appear in the conversation even when nobody is typing.
from libs import ScheduleService
schedules = ScheduleService(".cortana/schedules.db", {"reporter": agent})
await schedules.create(
id="daily-report",
agent_id="reporter",
cron="0 9 * * 1-5",
timezone="Africa/Johannesburg",
prompt="Prepare the morning report.",
)
await schedules.start()
Five-, six- and seven-field cron expressions and @hourly, @daily, @weekly, @monthly and @midnight are supported. Schedules can be paused, resumed, run now or deleted. The CLI stores them in the app's own .cortana/ directory whatever the workspace; set CORTANA_STATE_DIR to move it.