Coding agent

Muse Code

Meta's sandbox-first terminal agent with goals, memory, and multi-agent workflows.

Status
Active
License
Proprietary native binary
Evidence
Documented, 17 sources
Product record checked
2026-08-05

At a glance

Meta's terminal and CI coding agent with repository tools, default-on approvals and OS sandboxing, persistent memory, compaction, resumable sessions, goals, subagents, skills, hooks, and MCP.

Good choice if

  • Meta Model API users who want a Muse Spark agent for terminal and CI work
  • Long tasks using goals, compaction, persistent memory, and resumable sessions
  • Parallel workflows with subagents, observer agents, worktrees, skills, hooks, and MCP

Check before choosing

  • The beta ships as a proprietary binary, so evidence is documentation-verifiable without a public implementation audit
  • The Meta and Muse Spark path uses usage-based billing; model performance and benchmark results do not establish harness quality, and no built-in browser or non-Meta provider is documented
  • Default on-request approval passes commands outside a small dangerous set and relies on the sandbox; --yolo disables both controls and trusts repository instructions
See 3 more considerations
  • MCP tools run outside the filesystem and network sandbox, while hooks bypass both sandbox and approval
  • Worktree isolation is opt-in and falls back to a shared workspace outside Git; it does not contain processes or external effects
  • Resume rebuilds session state but not files; effects interrupted in flight retain an unknown outcome and require verification before retry, and no file checkpoint or rollback mechanism is documented; committed memory is injected before workspace trust

Capability support

Documented first-class product support, checked against the sources below.

External tools (MCP)
DocumentedProduct-supported MCP integrationSource · checked 2026-08-05The source establishes the mechanism, not its quality or availability in every mode.
Reusable skills
DocumentedProduct-supported reusable skill packagesSource · checked 2026-08-05Support does not establish portability, package quality, safety, or adoption.
Local models
Not documentedNo first-class support established by the current recordAbsence of current documentation is not proof that the capability is impossible.
Agent parallelism
DocumentedDelegated or parallel agent workflowSource · checked 2026-08-05The source establishes the mechanism, not its quality or availability in every mode.
Runs without an open UI
DocumentedNon-interactive or automation surfaceSource · checked 2026-08-05The source establishes the mechanism, not its quality or availability in every mode.
Browser control
Not documentedNo first-class support established by the current recordAbsence of current documentation is not proof that the capability is impossible.
Isolated execution
Available by defaultDefault OS-enforced shell sandbox with workspace-confined file tools and configurable network accessSource · checked 2026-08-05MCP tools are outside the OS sandbox, hooks bypass both sandbox and approval, and explicit flags can disable the guardrails.
Undo file changes
Not documentedNo first-class support established by the current recordAbsence of current documentation is not proof that the capability is impossible.

Getting started

Install `muse`, authenticate with Meta, review workspace trust and approvals, and keep the sandbox enabled outside an isolated environment.

Open official documentation

Classification and operating model

Category fit and technical mechanisms are evidence records, not product-quality scores.

Category fit
Qualifies, 4/4 criteria
Operating model
7/7 layers documented
Inspect category criteria and operating mechanismsFirst-party records

Why it qualifies as a coding harness

This confirms category fit, not product quality. Every required criterion links back to first-party evidence.

Qualifies4 of 4 required criteria evidenced

Membership establishes category fit only. It does not score quality, safety, autonomy, model capability, or benchmark performance. · Read the membership rule.

How it works under the hood

Seven mechanisms mapped from first-party records. These labels describe what the harness provides, not how intelligent its model is.

7/7layers documented
  • Execution & isolationSandbox availableDocumented mechanism, not a performance score.
  • Tooling & integrationsExtensible toolsDocumented mechanism, not a performance score.
  • Context & statePersistent stateDocumented mechanism, not a performance score.
  • Lifecycle & recoverySession resumeDocumented mechanism, not a performance score.
  • ObservabilityStructured tracesDocumented mechanism, not a performance score.
  • VerificationTool-assistedDocumented mechanism, not a performance score.
  • Governance & permissionsPolicy controlsDocumented mechanism, not a performance score.

Measured and public context

Configuration-specific measurements and source-native activity stay separate from general product capability.

Inspect code audit, measured configurations, and ecosystem signalsContext, not a product score

Public code audit

No official public repository was located for a code-level audit.

Measured configurations

No benchmark run passes the full metadata admission policy for this harness yet. Missing data is not scored as zero.

Benchmark policy and all runs

First-party evidence

Each capability claim links to the first-party record that supports it.

17 first-party sourcesProduct record checked

Product and interfaces

2 sources
View 1 more source

Execution and control

4 sources
View 3 more sources

Agents, state and recovery

6 sources
View 5 more sources

Automation and extensions

4 sources
View 3 more sources

Enterprise and operations

1 source