Coding agent

Factory Droid

Composable coding agent with autonomy tiers, worktrees, and custom workers.

Status
Active
License
Proprietary
Evidence
Documented, 17 sources
Product record checked
2026-08-12

At a glance

Factory's proprietary agent combines an interactive CLI, structured headless execution, IDE integrations, many hosted or local models, MCP, custom subagents, worktrees, and tiered permission levels from read-only analysis to deployment workflows.

Good choice if

  • Composable terminal and CI workflows that need structured streams and explicit autonomy tiers
  • Parallel changes isolated through named Git worktrees and custom delegated workers
  • Teams combining Factory-hosted models, BYOK providers, local Ollama, MCP, and IDE context

Check before choosing

  • The Private Preview OS sandbox is opt-in; its default per-command mode confines shell children while the main Droid process remains outside the OS boundary
  • Sandbox file reads allow all paths unless explicitly denied, allowed network domains remain exfiltration channels, and Factory recommends a container or VM for untrusted code
  • Git worktrees isolate file changes, not processes; dirty headless worktrees are preserved for review rather than silently removed
See 5 more considerations
  • Droid Control provides browser and desktop automation through a plugin rather than the minimal core tool set
  • Hooks run automatically with the current environment's credentials and may block tools or merge in plugin-provided code, so project and plugin hooks must be reviewed as trusted execution
  • Session rewind and file snapshots do not reverse external side effects, and persistent project knowledge relies on explicit AGENTS.md instructions rather than learned product memory
  • CLI session mirroring to Factory web is enabled by default and must be disabled with cloudSessionSync when local-only conversation storage is required
  • Factory publishes selected benchmark scorecards, but its public repository does not expose the closed agent implementation or a complete immutable evaluation corpus

Capability support

Documented first-class product support, checked against the sources below.

External tools (MCP)
DocumentedProduct-supported MCP integrationSource · checked 2026-07-30The source establishes the mechanism, not its quality or availability in every mode.
Reusable skills
Not documentedNo first-class support established by the current recordAbsence of current documentation is not proof that the capability is impossible.
Local models
DocumentedLocal or self-hosted model pathSource · checked 2026-07-30The source establishes the mechanism, not its quality or availability in every mode.
Agent parallelism
DocumentedDelegated or parallel agent workflowSource · checked 2026-07-30The source establishes the mechanism, not its quality or availability in every mode.
Runs without an open UI
DocumentedNon-interactive or automation surfaceSource · checked 2026-07-30The source establishes the mechanism, not its quality or availability in every mode.
Browser control
DocumentedBuilt-in or product-supported browser controlSource · checked 2026-07-30The source establishes the mechanism, not its quality or availability in every mode.
Isolated execution
OptionalPrivate Preview OS sandbox; opt-inSource · checked 2026-07-30Per-command mode contains shell children while the main Droid process remains outside the boundary.
Undo file changes
DocumentedInteractive session rewindSource · checked 2026-07-30Rewind does not reverse external side effects.

Getting started

Install `droid`, authenticate with Factory or configure BYOK, then begin in read-only mode and raise autonomy only inside a reviewed environment.

Open official documentation

Classification and operating model

Category fit and technical mechanisms are evidence records, not product-quality scores.

Category fit
Qualifies, 4/4 criteria
Operating model
7/7 layers documented
Inspect category criteria and operating mechanismsFirst-party records

Why it qualifies as a coding harness

This confirms category fit, not product quality. Every required criterion links back to first-party evidence.

Qualifies4 of 4 required criteria evidenced
  • Adaptive agent loop

    Documented

    The system repeatedly observes results and chooses the next action instead of following a fixed one-pass graph.

  • Repository tool execution

    Documented

    The system can use tools to inspect and change a repository or its execution environment.

  • Task-aware context management

    Documented

    The runtime assembles, updates, compacts, retrieves, or persists task-relevant context while work proceeds.

  • Model-independent runtime control

    Documented

    Permissions, budgets, interruption, policy, or stop controls operate outside the model's own text generation.

Membership establishes category fit only. It does not score quality, safety, autonomy, model capability, or benchmark performance. · Read the membership rule.

How it works under the hood

Seven mechanisms mapped from first-party records. These labels describe what the harness provides, not how intelligent its model is.

7/7layers documented
  • Execution & isolationSandbox availableDocumented mechanism, not a performance score.
  • Tooling & integrationsExtensible + browserDocumented mechanism, not a performance score.
  • Context & stateManaged contextDocumented mechanism, not a performance score.
  • Lifecycle & recoveryCheckpoint/rewindDocumented mechanism, not a performance score.
  • ObservabilityStructured tracesDocumented mechanism, not a performance score.
  • VerificationTool-assistedDocumented mechanism, not a performance score.
  • Governance & permissionsPolicy controlsDocumented mechanism, not a performance score.

Measured and public context

Configuration-specific measurements and source-native activity stay separate from general product capability.

Inspect code audit, measured configurations, and ecosystem signalsContext, not a product score

Public code audit

Unrankedsupport-only repository
Security policy
Not found
CI workflow
Present at inspected commit
Automated tests
Not found
Evaluation assets
Present at inspected commit
Contributor documentation
Not found

The 440-blob public support tree was inspected at README.md, .github/workflows/, docs/benchmarks/, and docs/snippets/leaderboards/. It contains product documentation, selected vendor benchmark pages, and six workflows, but no Droid implementation, engineering tests, repository security policy, contributor guide, complete eval corpus, or immutable result record. The root README reserves all rights, and published scorecards remain separate from product classification.

Inspect commit 1fd9026d72f81668d88f37237cb5a2e89a17e6e2, checked 2026-08-12

Measured configurations

No benchmark run passes the full metadata admission policy for this harness yet. Missing data is not scored as zero.

Benchmark policy and all runs
Context, not quality

Public ecosystem signals

Source-native observations for exact mapped artifacts and reviewed stable release trains. Different units and populations stay separate, and missing coverage is never treated as zero.

View this harness in Usage

Routing, package retrievals, release downloads, editor installs, and repository interest observe different populations. They are never added together and never affect capability evidence, classification, or measured results.

Interpretation rulesSignals checked

First-party evidence

Each capability claim links to the first-party record that supports it.

17 first-party sourcesProduct record checked

Product and interfaces

4 sources
View 3 more sources

Execution and control

3 sources
View 2 more sources

Agents, state and recovery

2 sources
View 1 more source

Automation and extensions

3 sources
View 2 more sources

Enterprise and operations

2 sources
View 1 more source

Releases and public code audit

3 sources
View 2 more sources