Extensible harness

Codebuff

Autonomous multi-agent coding from the terminal or TypeScript.

Status
Active
License
Apache-2.0 code with hosted service
Evidence
Documented + code-verifiable, 16 sources
Product record checked
2026-07-27

At a glance

An open-source coding harness that coordinates specialized agents for finding files, editing, review, tests, research, and browser work. Use it interactively in the terminal or build deterministic agent workflows with its TypeScript SDK.

Good choice if

  • Developers who want the agent to complete terminal tasks with few interruptions
  • Teams that want file discovery, implementation, review, tests, research, and browser work coordinated by specialized agents
  • Programmatic coding pipelines that need provider choice and deterministic TypeScript branching

Check before choosing

  • It runs edits and commands on the host without per-command permission prompts by default
  • Isolation requires the provided Docker setup and a scoped code copy; no built-in sandbox is documented
  • Large-repository support is vendor-documented through code maps and VS Code-scale testing, but no independent scale result or service level is published
See 1 more considerations
  • Reliable custom workflows require writing and maintaining TypeScript agent definitions

Why it qualifies as a coding harness

This confirms category fit, not product quality. Every required criterion links back to first-party evidence.

Qualifies4 of 4 required criteria evidenced
  • Adaptive agent loop

    Documented

    The system repeatedly observes results and chooses the next action instead of following a fixed one-pass graph.

  • Repository tool execution

    Documented

    The system can use tools to inspect and change a repository or its execution environment.

  • Task-aware context management

    Documented

    The runtime assembles, updates, compacts, retrieves, or persists task-relevant context while work proceeds.

  • Model-independent runtime control

    Documented

    Permissions, budgets, interruption, policy, or stop controls operate outside the model's own text generation.

Membership establishes category fit only. It does not score quality, safety, autonomy, model capability, or benchmark performance. · Read the membership rule.

How it works under the hood

Seven mechanisms mapped from first-party records. These labels describe what the harness provides, not how intelligent its model is.

7/7layers documented
  • Execution & isolationHost processDocumented mechanism, not a performance score.
  • Tooling & integrationsExtensible + browserDocumented mechanism, not a performance score.
  • Context & stateManaged contextDocumented mechanism, not a performance score.
  • Lifecycle & recoverySession resumeDocumented mechanism, not a performance score.
  • ObservabilityLogs/transcriptsDocumented mechanism, not a performance score.
  • VerificationTool-assistedDocumented mechanism, not a performance score.
  • Governance & permissionsHost accessDocumented mechanism, not a performance score.

Public code audit

4/5public artifacts present
Security policy
Present at inspected commit
CI workflow
Not found
Automated tests
Present at inspected commit
Evaluation assets
Present at inspected commit
Contributor documentation
Present at inspected commit

The current public snapshot contains 363 test-like files, security and contributor documentation, and 38 BuffBench assets, but no public CI configuration or immutable result record. BuffBench is project-owned and LLM-judged; its README describes two parallel judges while its diagram still says three, so no score is imported.

Inspect commit 180071751c43a479684672576c44f14e120d2717, checked 2026-07-28

Measured configurations

No benchmark run passes the full metadata admission policy for this harness yet. Missing data is not scored as zero.

Benchmark policy and all runs

Capability support

Documented first-class product support, checked against the sources below.

External tools (MCP)
DocumentedProduct-supported MCP integrationSource · checked 2026-07-27The source establishes the mechanism, not its quality or availability in every mode.
Local models
Not documentedNo first-class support established by the current recordAbsence of current documentation is not proof that the capability is impossible.
Agent parallelism
DocumentedDelegated or parallel agent workflowSource · checked 2026-07-27The source establishes the mechanism, not its quality or availability in every mode.
Runs without an open UI
DocumentedNon-interactive or automation surfaceSource · checked 2026-07-27The source establishes the mechanism, not its quality or availability in every mode.
Browser control
DocumentedBuilt-in or product-supported browser controlSource · checked 2026-07-27The source establishes the mechanism, not its quality or availability in every mode.
Isolated execution
Not documentedNo first-class support established by the current recordAbsence of current documentation is not proof that the capability is impossible.
Undo file changes
Not documentedNo first-class support established by the current recordAbsence of current documentation is not proof that the capability is impossible.

Primary evidence

Each capability claim is tied to a first-party record and a verification date.

Product record checked 2026-07-27
View 8 additional sources

Ecosystem discovery

OpenRouter coding apps Discovery signal only; usage rank is not used as a quality or capability score. Observed 2026-07-27.