Quickstart

Get Muse Code running in minutes.

Muse Code is generally available for macOS and Linux, and runs on Windows through WSL2. It takes on complex software engineering tasks across large repositories — planning changes, writing code, and validating the results — entirely from your terminal.

1. Install

One command installs the CLI:

$curl -fsSL https://dev.meta.ai/install.sh | bash

The script drops the muse binary into your PATH — verify with muse --version. Prefer to review the script first? Download it with -o install.sh, audit, then run.

Windows (WSL2)

There is no native Windows build — the installer supports macOS and Linux. On Windows 10 or 11, run wsl --install in an administrator PowerShell to set up WSL2 with Ubuntu, open the Linux shell, and run the same one-line installer there. Keep your repositories on the Linux filesystem (for example ~/code) rather than under /mnt/c: file access across the boundary is much slower.

2. Authenticate

Your first run opens the browser to authenticate at dev.meta.ai with your Meta developer account. Credentials cache locally — no re-auth per session. Usage bills to your monthly plan or, on pay-as-you-go, to the Meta Model API.

Non-interactive runs. For CI, scripts, and headless machines, use an API key instead of the browser flow: create one in the Model API dashboard (API keys → Create API key) and set it as MODEL_API_KEY.

3. Run your first task

Once auth is done, you’re free to get started in any of your project directories:

$ cd my-repo$ muse

Describe the outcome you want in plain language. A good first task is small and verifiable — “add a missing test for parseDate and make sure it passes” — you’ll see a worker draft the change, a reviewer critique it, and the test run.

Two cost levers worth knowing from day one: Muse Spark is a reasoning model — it thinks before it answers, and those tokens bill as output — so /effort dials the thinking up or down per task. And your tone, format, and other standing rules can live in a system message so you don’t repeat them in every prompt.

Async background agents

Muse Code operates with a simple agent loop plus a set of async background agents that enhance the main agent’s capability. These specialized agents remain active throughout each session rather than being spawned per task, which avoids redundant information gathering. They carry out next steps on their own and choose when to communicate back to the main agent — their persistence reduces latency and the need for steering on difficult, multi-step tasks.

Agent fan-out & git worktrees

When a job splits into several tasks, they fan out automatically to separate agents: the parent spawns a write-capable child per task, and each child gets its own git worktree — so parallel children never collide on the same files and your working copy stays clean.

Workflows

Added when Muse Code left beta, Workflows orchestrate a team of specialized agents across a larger job: each stage hands its intermediate work to the next, and you get a single result at the end. Start one by asking Muse Code to “use a workflow”; setting effort to ultra can trigger one automatically. /workflows shows every run live — stop, restart, or cancel it there — and saved workflows can be reused.

Inter-session messaging

Sessions running on the same machine can message each other over a local Unix socket. When a change in one session affects what another is building, it passes a warning across; when one session settles a question another is blocked on, it sends the answer — no retyping between terminals. The agent discovers its peers and sends messages with two built-in tools.

Event log & rewind

Muse Code uses a local event log to which every model call, tool run, approval, and edit is appended. This single source of truth makes the runtime replay-exact and restart-safe: after a crash, the agent resumes precisely where it stopped, letting it take on long-running tasks without being derailed by failures. Every subagent spawn, steer, and cancel is observable — muse replay walks a session step by step, and full traces export for handoff or audit. Press Esc twice to rewind the conversation to an earlier point; Muse Code only offers rewind points the event log marks as safe.

Bundled skills & commands

Muse Code ships with several default skills — put a rough idea in, get a grilled, taste-checked feature out:

/planTurns a task into an approval-gated plan before any code changes
/grillStress-tests that plan until it holds up
/goalWorks toward successful completion of the specified objective — what you originally wanted is what gets merged
/workflowsLists workflow runs live — stop, restart, or cancel any of them
/modelSwitches the backing model — e.g. to muse-spark-1.3 at standard Model API pricing
/effortDials reasoning up or down — thinking tokens bill as output, so save deep reasoning for hard tasks
Esc EscRewinds the conversation to a safe point from the event log
muse replayWalks a session's event log step by step; full traces export for handoff or audit

Muse Spark 1.3

Muse Spark 1.3, released September 2, 2026, is the model behind Muse Code. It keeps the training recipe that set 1.2 apart — co-training with the Muse Code harness, extensive long-horizon work, and a self-improvement loop — and focuses on real-world usability. Pricing is unchanged from 1.2. Three things stand out:

Longer threads
Sustains longer-horizon work across multiple workflows in a single thread, carrying context from one stage to the next.
Asks instead of guessing
Collaborates more actively: it asks clarifying questions, flags when it’s stuck, and is better calibrated about its own limits instead of reporting outcomes it didn’t reach.
Leaner runs
In Meta engineers’ comparisons it finished the same coding work with about 20% fewer tool calls and 25% fewer tokens than 1.2 — at the same per-token price, that is a smaller bill.

Benchmarks (max reasoning configuration, per Meta): DeepSWE 1.1 · 75.4% — Terminal-Bench 2.1 · 88.8% — SWEAtlas CodeBase QnA · 59.4%. At launch the max configuration was still in safety testing; the configuration most developers get scores lower. Methodology at research.meta.ai.

Case study: kernel optimization

Meta tested Muse Spark 1.2’s ability to iteratively optimize GPU kernels over 1,000+ tool calls — up to 24 hours. Inside Muse Code’s agentic environment, the model writes, compiles, profiles, and progressively improves kernel performance against a baseline, benchmarked on KDA and MLA kernels for NVIDIA Hopper GPUs. Barred from wrapping third-party kernel libraries, it implemented the algorithms in Triton directly — and kept finding substantial improvements over the baseline.

Meta Model API

The same models behind Muse Code are available directly — Muse Spark 1.3 ships in Muse Code, the Meta Model API, OpenRouter, and Cursor. The API is drop-in compatible with the OpenAI SDK, the Anthropic SDK, and agent CLIs like OpenCode and Claude Code: point your client at the base URL and keep the rest of your code. The context window is 1,048,576 tokens.

# export MODEL_API_KEY="LLM|{numeric_id}|{secret}"client = OpenAI( base_url="https://api.meta.ai/v1", api_key=os.environ["MODEL_API_KEY"], ) response = client.chat.completions.create( model="muse-spark-1.3", messages=[{"role": "user", "content": "Hello, world!"}], )

Muse Code SDK

In developer preview since Muse Code left beta, the SDK lets teams build their own agents on the Muse Session Protocol (MSP) — an open protocol over standard I/O with no server and no network, so everything stays on your machine. The TypeScript library covers sessions, tools, and permission control, and MSP’s machine-readable schema lets you generate clients in other languages. Source: github.com/meta-models/muse-code-sdk.

Cookbook

Recipes that run the first time you copy them — each one solves a focused problem, shows working code, and points to what’s next. Three sections, plus recipes for Muse Code’s core patterns.Browse all 28 recipes →

API fundamentals · 10 recipes
Prove each primitive works: chat completions, streaming, tool calling, structured output, prompt caching, reasoning tokens, vision input, long context, error handling, and search grounding with inline citations.
Agent patterns · 5 recipes
The loops that turn a model into an agent: the perceive-decide-act loop, interleaved reasoning and tool use, multi-turn context management, and validated in-place edits.
Use cases · 13 recipes
End to end: browser-verified web design and game dev, a four-profile multi-agent product studio coordinating through a shared Kanban board, an autonomous GitHub Actions bot on OpenCode, computer use on Linux and macOS, and more.
Muse Code recipes
Agent fan-out — split one big job across subagents in isolated worktrees; bundled skills — a rough idea in, a grilled, taste-checked feature out; goal tracking — keep an agent on task until it merges what you wanted.