Choose the right coding harness.

The model reasons. The harness is the CLI, IDE extension, or agent platform that turns it into working code.

Active catalog entries
35
Latest source check
Jul 28, 2026
Peer-reviewed studies
21
Measured configurations
5

Choose a workflow

Start with a familiar setup. See the strongest match and what to check before choosing.

Describe the outcome in an IDE, let the agent write most of the code, and review the important steps.

Interface
IDE
Model access
Subscription
Control
Review risky actions
Start hereClaude Codestrong workflow match. Top three in 97% of priority variations.
Why it fits

Documented use case: Developers already paying for Claude Pro or Max.

Check before choosing

The local CLI is host-first: OS sandboxing is disabled by default, may fall back to unsandboxed execution when unavailable unless failIfUnavailable is enabled, and covers Bash subprocesses rather than every tool

Claude Code is the strongest match for this workflow and stays in the top three in 97 percent of tested priority variations.

Ranking stability

How often each harness stays in the top three across 512 priority variations. This is not task success.

Order follows the reference fit
  1. Claude Codestrong match, evidence: Documented + independently measured configuration97%strong match. Top three in 97 percent of 512 tested priority variations, with rank range 1 to 7. Catalog membership and every must-have have current supporting documentation. Evidence state: Documented + independently measured configuration. Opens the evidence profile.
  2. Aiderstrong match, evidence: Documented + code-verifiable79%strong match. Top three in 79 percent of 512 tested priority variations, with rank range 1 to 10. Catalog membership and every must-have have current supporting documentation. Evidence state: Documented + code-verifiable. Opens the evidence profile.
  3. Grok Buildstrong match, evidence: Documented + code-verifiable35%strong match. Top three in 35 percent of 512 tested priority variations, with rank range 1 to 12. Catalog membership and every must-have have current supporting documentation. Evidence state: Documented + code-verifiable. Opens the evidence profile.
  4. Command Codestrong match, evidence: Documented31%strong match. Top three in 31 percent of 512 tested priority variations, with rank range 1 to 13. Catalog membership and every must-have have current supporting documentation. Evidence state: Documented. Opens the evidence profile.
  5. Clinestrong match, evidence: Documented + code-verifiable17%strong match. Top three in 17 percent of 512 tested priority variations, with rank range 1 to 17. Catalog membership and every must-have have current supporting documentation. Evidence state: Documented + code-verifiable. Opens the evidence profile.
  6. Mistral Vibestrong match, evidence: Documented + code-verifiable17%strong match. Top three in 17 percent of 512 tested priority variations, with rank range 1 to 16. Catalog membership and every must-have have current supporting documentation. Evidence state: Documented + code-verifiable. Opens the evidence profile.
  7. Qwen Codestrong match, evidence: Documented + code-verifiable8%strong match. Top three in 8 percent of 512 tested priority variations, with rank range 1 to 18. Catalog membership and every must-have have current supporting documentation. Evidence state: Documented + code-verifiable. Opens the evidence profile.
View all 7 assumptionsPriority, scope, work mode, and must-haves
Interface
IDE
Priority
Simplicity
Model access
Subscription
Control
Review risky actions
Change scope
Cross-file
Operating mode
Work together
Required
None
How this ranking worksOrdering, stability, and deal-breakers

Why this orderTools that fit your priorities appear first.

StabilityHow often a tool stays near the top when your priorities change slightly.

Deal-breakersNeighboring product layers and tools without current evidence for one are left out instead of receiving a lower score.

We first require documented coding-harness membership, then remove tools without current documentation for a must-have. The remaining products are ordered using published reference weights: main priority 30, approval style 25, change size 25, and work mode 20. We then vary those priorities 512 ways. The percentage is a stability check, not a task success rate.

See all 22 eligible harnesses

22 of 35 active catalog entries pass coding-harness membership and every must-have in this workflow.

  1. 1Claude CodeDocumented use case: Developers already paying for Claude Pro or Max.Top 3 in 97%strong match, average position 1.4, Documented + independently measured configuration
  2. 2AiderDocumented use case: Developers who prefer a tight Git-centric editing loop.Top 3 in 79%strong match, average position 2.6, Documented + code-verifiable
  3. 3Grok BuildDocumented use case: Plan-review-execute workflows with explicit diffs and permission controls.Top 3 in 35%strong match, average position 4.8, Documented + code-verifiable
  4. 4Command CodeDocumented use case: Terminal users who need explicit allow, ask, and deny policy rules.Top 3 in 31%strong match, average position 5, Documented
  5. 5ClineDocumented use case: IDE users who want to approve actions and inspect diffs.Top 3 in 17%strong match, average position 6.6, Documented + code-verifiable
  6. 6Mistral VibeDocumented use case: Vibe coders who want approvals, a plan mode, and file rewind without giving up a capable agent.Top 3 in 17%strong match, average position 6.8, Documented + code-verifiable
  7. 7Qwen CodeDocumented use case: Terminal and CI workflows that can explicitly enable a first-party sandbox path.Top 3 in 8%strong match, average position 8.5, Documented + code-verifiable
  8. 8ZCodeDocumented use case: Developers who prefer a visual agent workspace over a terminal-only loop.Top 3 in 6%strong match, average position 8.7, Documented
  9. 9stagewiseDocumented use case: Vibe coders who want the browser, DOM context, app preview, and code changes in one desktop workspace.Top 3 in 4%strong match, average position 9.9, Documented + code-verifiable
  10. 10Junie CLIDocumented use case: JetBrains users who want one agent across an IDE, terminal, and CI.Top 3 in 3%strong match, average position 10.2, Documented
  11. 11CodexDocumented use case: ChatGPT subscribers who want one agent across app, IDE, and terminal.Top 3 in 1%strong match, average position 10.7, Documented + code-verifiable + independently measured configuration
  12. 12OpenCodeDocumented use case: Developers who switch between providers and subscription paths.Top 3 in 1%strong match, average position 10.9, Documented + code-verifiable
  13. 13Kilo CodeDocumented use case: Vibe coders who want an IDE-first agent, subscription or local models, and a practical undo path.Top 3 in 0%strong match, average position 12.6, Documented + code-verifiable
  14. 14Kimi CodeDocumented use case: Parallel coding work that uses foreground, background, or swarm-style subagents.Top 3 in 0%strong match, average position 14.1, Documented + code-verifiable
  15. 15Poolside Agent CLIDocumented use case: Developers who want one TUI to drive Poolside, OpenRouter, Ollama, local OpenAI-compatible models, or other ACP agents.Top 3 in 0%strong match, average position 14.5, Documented
  16. 16Zoo CodeDocumented use case: Vibe coders who want proposed tools and diffs visible in VS Code before they run.Top 3 in 0%strong match, average position 13.3, Documented + code-verifiable
  17. 17Factory DroidDocumented use case: Composable terminal and CI workflows that need structured streams and explicit autonomy tiers.Top 3 in 0%strong match, average position 16.6, Documented
  18. 18Hermes AgentDocumented use case: Long-lived development workflows that benefit from bounded memory, searchable sessions, and persistent goals.Top 3 in 0%strong match, average position 16.6, Documented + code-verifiable
  19. 19AmpDocumented use case: Hands-off vibe coding in disposable managed cloud orbs.Top 3 in 0%strong match, average position 17.8, Documented
  20. 20MuxDocumented use case: Parallel feature work in isolated child workspaces.Top 3 in 0%good match, average position 19.1, Documented + code-verifiable
  21. 21Oh My PiDocumented use case: Terminal users who want IDE-grade code intelligence and debugger access.Top 3 in 0%good match, average position 20.5, Documented + code-verifiable
  22. 22OpenHandsDocumented use case: Longer autonomous tasks in an isolated runtime.Top 3 in 0%good match, average position 22, Documented + code-verifiable
Why 13 harnesses do not match

These products are not ranked because they are outside the default coding-harness layer or at least one required condition is not currently documented.

Highlights

Three separate evidence views. Select one to inspect the complete ranking without turning them into an overall product score.

Definitions and weights

Seven source-backed architecture layers. Select one layer at a time: ordinal mechanisms are never summed into a universal product grade.

Hermes Agent is first in the selected operational mechanisms view: Workflow-gated.

Show the complete 35-record ranking
  1. 1Hermes Agent
  2. 2ZCode
  3. 3Aider
  4. 4Amp
  5. 5Antigravity CLI
  6. 6Claude Code
  7. 7Cline
  8. 8Codebuff
  9. 9Coder Agents
  10. 10Codex
  11. 11Command Code
  12. 12Crush
  13. 13Cursor CLI
  14. 14Factory Droid
  15. 15ForgeCode
  16. 16Gemini CLI
  17. 17GitHub Copilot CLI
  18. 18goose
  19. 19Grok Build
  20. 20Junie CLI
  21. 21Kilo Code
  22. 22Kimi Code
  23. 23Kiro CLI
  24. 24Letta Harness
  25. 25Mistral Vibe
  26. 26Mux
  27. 27Oh My Pi
  28. 28OpenCode
  29. 29OpenHands
  30. 30Poolside Agent CLI
  31. 31Qwen Code
  32. 32stagewise
  33. 33Zoo Code
  34. 34mini-SWE-agent
  35. 35Pi Agent

35 comparable records. Missing evidence is never renormalized into a better result.

Three rules behind the ranking.

The short version of the research. Every rule links to its underlying papers.

Open the full methodology

Browse all harnesses.

Filter by a capability, then open a profile for trade-offs and sources.

35 active profiles

Advanced filtersLayer, role, surface, and runtime
Coding harness56 first-party sources

Polished Claude-first agent across terminal, IDE, desktop, and web.

Role
Coding agent
Interfaces
Terminal, IDE, Web / desktop, Automation
Model access
Enterprise routing
Checked 2026-07-27View profile
Coding harness46 first-party sources

Agentic coding across app, terminal, IDE, cloud, and automation.

Role
Coding agent
Interfaces
Terminal, IDE, Web / desktop, Automation
Model access
Multi-provider
Checked 2026-07-27View profile
Coding harness25 first-party sources

Model-agnostic coding agent with a broad, configurable surface.

Role
Coding agent
Interfaces
Terminal, IDE, Web / desktop, Automation
Model access
Multi-provider
Checked 2026-07-27View profile
Coding harness19 first-party sources

Minimal terminal harness designed to be extended, embedded, or automated.

Role
Extensible harness
Interfaces
Terminal, Automation
Model access
Multi-provider
Checked 2026-07-27View profile
Coding harness18 first-party sources

A batteries-included Pi fork with LSP, debugger, browser, and subagents.

Role
Coding agent
Interfaces
Terminal, IDE, Automation
Model access
Multi-provider
Checked 2026-07-28View profile
Coding harness24 first-party sources

Extensible Rust coding agent with plan review, sandbox profiles, and ACP.

Role
Coding agent
Interfaces
Terminal, IDE, Automation
Model access
Multi-provider
Checked 2026-07-27View profile
Coding harness18 first-party sources

Focused pair programming built around Git, diffs, and explicit control.

Role
Pair programmer
Interfaces
Terminal, IDE, Web / desktop, Automation
Model access
Multi-provider
Checked 2026-07-27View profile
Coding harness23 first-party sources

Agent platform with selectable local, sandboxed, and remote runtimes.

Role
Agent platform
Interfaces
Web / desktop, Terminal, IDE, Automation
Model access
Multi-provider
Checked 2026-07-27View profile

Capability filters reflect explicit product documentation. They do not compare model intelligence or benchmark performance.

Read the methodology