Compare measured configurations.

Results belong to an exact model, harness version, environment, budget, and run policy.

Terminal-Bench 2.1

Accuracy after integrity review. 89 tasks, five attempts per task, 445 trials per configuration.

Official leaderboard
  1. Claude Codeanthropic/claude-fable-5, harness 2.1.16783.82

    95% descriptive interval 81.5-86.1

  2. Codexopenai/gpt-5.5, harness 0.125.083.15

    95% descriptive interval 80.9-85.4

  3. Cursor CLIcursor/grok-4.5, harness 2026.07.08-0c04a8a79.33

    95% descriptive interval 76.5-82.2

  4. mini-SWE-agentopenai/muse-spark-1.1, harness 2.4.576.18

    95% descriptive interval 73.8-78.6

  5. Gemini CLIgemini/gemini-3-pro-preview, harness 0.40.065.84

    95% descriptive interval 63.1-68.5

Reproducibility records

Open a record to inspect budget, environment, token use, source, and integrity handling.

Claude Code 2.1.16783.82%, 95% interval 81.5-86.1
Model
anthropic/claude-fable-5
Reasoning effort
xhigh
Benchmark
Terminal-Bench 2.1
Attempts
89 tasks × 5 = 445 trials
Cost
$552.67
Cost / observed success
$1.48
Accuracy/cost frontier
Non-dominated
Top interval group
Overlaps leading 95% interval
Average duration
579.7 seconds per trial
Tokens
20,480,925 input, 174,066,133 cached, 9,954,945 output
Integrity review
1 disqualified trials, 0.22% adjustment
Run date
2026-06-07
Verified
2026-07-27

Environment: Harbor-managed Docker task containers using the official Terminal-Bench 2.1 dataset.

Network: Task-defined network access; passing trajectories were reviewed for solution lookup and reward hacking.

Budget: Official per-task time and resource limits with no timeout or resource overrides; five trials per task.

Codex 0.125.083.15%, 95% interval 80.9-85.4
Model
openai/gpt-5.5
Reasoning effort
xhigh
Benchmark
Terminal-Bench 2.1
Attempts
89 tasks × 5 = 445 trials
Cost
$2,059.19
Cost / observed success
$5.57
Accuracy/cost frontier
Dominated by another admitted system
Top interval group
Overlaps leading 95% interval
Average duration
482.6 seconds per trial
Tokens
336,797,311 input, 392,433,664 cached, 5,966,373 output
Integrity review
1 disqualified trials, 0.22% adjustment
Run date
2026-05-01
Verified
2026-07-27

Environment: Harbor-managed Docker task containers using the official Terminal-Bench 2.1 dataset.

Network: Task-defined network access; passing trajectories were reviewed for solution lookup and reward hacking.

Budget: Official per-task time and resource limits with no timeout or resource overrides; five trials per task.

Cursor CLI 2026.07.08-0c04a8a79.33%, 95% interval 76.5-82.2
Model
cursor/grok-4.5
Reasoning effort
high
Benchmark
Terminal-Bench 2.1
Attempts
89 tasks × 5 = 445 trials
Cost
$134.09
Cost / observed success
$0.38
Accuracy/cost frontier
Non-dominated
Top interval group
Overlaps leading 95% interval
Average duration
443.0 seconds per trial
Tokens
12,570,445 input, 161,079,936 cached, 4,734,062 output
Integrity review
40 disqualified trials, 8.99% adjustment
Run date
2026-07-09
Verified
2026-07-27

Environment: Harbor-managed Docker task containers using the official Terminal-Bench 2.1 dataset.

Network: Task-defined network access; passing trajectories were reviewed for solution lookup and reward hacking.

Budget: Official per-task time and resource limits with no timeout or resource overrides; five trials per task.

mini-SWE-agent 2.4.576.18%, 95% interval 73.8-78.6
Model
openai/muse-spark-1.1
Reasoning effort
xhigh
Benchmark
Terminal-Bench 2.1
Attempts
89 tasks × 5 = 445 trials
Cost
$198.05
Cost / observed success
$0.58
Accuracy/cost frontier
Dominated by another admitted system
Top interval group
Does not overlap leading interval
Average duration
687.0 seconds per trial
Tokens
60,996 input, 932,239,373 cached, 13,679,953 output
Integrity review
0 disqualified trials, 0.00% adjustment
Run date
2026-07-09
Verified
2026-07-27

Environment: Harbor-managed Docker task containers using the official Terminal-Bench 2.1 dataset.

Network: Task-defined network access; passing trajectories were reviewed for solution lookup and reward hacking.

Budget: Official per-task time and resource limits with no timeout or resource overrides; five trials per task.

Gemini CLI 0.40.065.84%, 95% interval 63.1-68.5
Model
gemini/gemini-3-pro-preview
Reasoning effort
high
Benchmark
Terminal-Bench 2.1
Attempts
89 tasks × 5 = 445 trials
Cost
$247.76
Cost / observed success
$0.85
Accuracy/cost frontier
Dominated by another admitted system
Top interval group
Does not overlap leading interval
Average duration
436.6 seconds per trial
Tokens
47,653,156 input, 329,123,284 cached, 7,219,388 output
Integrity review
2 disqualified trials, 0.45% adjustment
Run date
2026-05-01
Verified
2026-07-27

Environment: Harbor-managed Docker task containers using the official Terminal-Bench 2.1 dataset.

Network: Task-defined network access; passing trajectories were reviewed for solution lookup and reward hacking.

Budget: Official per-task time and resource limits with no timeout or resource overrides; five trials per task.

Admission fields

  • Exact harness version and model
  • Benchmark and dataset version
  • Reasoning effort and token usage
  • Sandbox, network, and budget policy
  • Tasks, attempts, and total trials
  • Cost, duration, run date, and primary source
  • Integrity review and disqualified trials

Claim boundary

The same harness can move when the model, effort, version, task set, budget, or integrity policy changes.

result = model × harness × config × environment × budget

HarnessMatch never copies this accuracy into workflow fit, operational readiness, or code auditability.

The displayed 95% interval is a normal approximation from the reported standard error. It is descriptive, not a task-cluster-corrected equivalence test.

Research behind the policy

These papers define reproducibility, episode traces, cost-aware evaluation, versioning, and contamination controls.