Compare measured configurations.
Results belong to an exact model, harness version, environment, budget, and run policy.
Terminal-Bench 2.1
Accuracy after integrity review. 89 tasks, five attempts per task, 445 trials per configuration.
Reproducibility records
Open a record to inspect budget, environment, token use, source, and integrity handling.
Claude Code 2.1.16783.82%, 95% interval 81.5-86.1
- Model
- anthropic/claude-fable-5
- Reasoning effort
- xhigh
- Benchmark
- Terminal-Bench 2.1
- Attempts
- 89 tasks × 5 = 445 trials
- Cost
- $552.67
- Cost / observed success
- $1.48
- Accuracy/cost frontier
- Non-dominated
- Top interval group
- Overlaps leading 95% interval
- Average duration
- 579.7 seconds per trial
- Tokens
- 20,480,925 input, 174,066,133 cached, 9,954,945 output
- Integrity review
- 1 disqualified trials, 0.22% adjustment
- Run date
- 2026-06-07
- Verified
- 2026-07-27
Environment: Harbor-managed Docker task containers using the official Terminal-Bench 2.1 dataset.
Network: Task-defined network access; passing trajectories were reviewed for solution lookup and reward hacking.
Budget: Official per-task time and resource limits with no timeout or resource overrides; five trials per task.
Codex 0.125.083.15%, 95% interval 80.9-85.4
- Model
- openai/gpt-5.5
- Reasoning effort
- xhigh
- Benchmark
- Terminal-Bench 2.1
- Attempts
- 89 tasks × 5 = 445 trials
- Cost
- $2,059.19
- Cost / observed success
- $5.57
- Accuracy/cost frontier
- Dominated by another admitted system
- Top interval group
- Overlaps leading 95% interval
- Average duration
- 482.6 seconds per trial
- Tokens
- 336,797,311 input, 392,433,664 cached, 5,966,373 output
- Integrity review
- 1 disqualified trials, 0.22% adjustment
- Run date
- 2026-05-01
- Verified
- 2026-07-27
Environment: Harbor-managed Docker task containers using the official Terminal-Bench 2.1 dataset.
Network: Task-defined network access; passing trajectories were reviewed for solution lookup and reward hacking.
Budget: Official per-task time and resource limits with no timeout or resource overrides; five trials per task.
Cursor CLI 2026.07.08-0c04a8a79.33%, 95% interval 76.5-82.2
- Model
- cursor/grok-4.5
- Reasoning effort
- high
- Benchmark
- Terminal-Bench 2.1
- Attempts
- 89 tasks × 5 = 445 trials
- Cost
- $134.09
- Cost / observed success
- $0.38
- Accuracy/cost frontier
- Non-dominated
- Top interval group
- Overlaps leading 95% interval
- Average duration
- 443.0 seconds per trial
- Tokens
- 12,570,445 input, 161,079,936 cached, 4,734,062 output
- Integrity review
- 40 disqualified trials, 8.99% adjustment
- Run date
- 2026-07-09
- Verified
- 2026-07-27
Environment: Harbor-managed Docker task containers using the official Terminal-Bench 2.1 dataset.
Network: Task-defined network access; passing trajectories were reviewed for solution lookup and reward hacking.
Budget: Official per-task time and resource limits with no timeout or resource overrides; five trials per task.
mini-SWE-agent 2.4.576.18%, 95% interval 73.8-78.6
- Model
- openai/muse-spark-1.1
- Reasoning effort
- xhigh
- Benchmark
- Terminal-Bench 2.1
- Attempts
- 89 tasks × 5 = 445 trials
- Cost
- $198.05
- Cost / observed success
- $0.58
- Accuracy/cost frontier
- Dominated by another admitted system
- Top interval group
- Does not overlap leading interval
- Average duration
- 687.0 seconds per trial
- Tokens
- 60,996 input, 932,239,373 cached, 13,679,953 output
- Integrity review
- 0 disqualified trials, 0.00% adjustment
- Run date
- 2026-07-09
- Verified
- 2026-07-27
Environment: Harbor-managed Docker task containers using the official Terminal-Bench 2.1 dataset.
Network: Task-defined network access; passing trajectories were reviewed for solution lookup and reward hacking.
Budget: Official per-task time and resource limits with no timeout or resource overrides; five trials per task.
Gemini CLI 0.40.065.84%, 95% interval 63.1-68.5
- Model
- gemini/gemini-3-pro-preview
- Reasoning effort
- high
- Benchmark
- Terminal-Bench 2.1
- Attempts
- 89 tasks × 5 = 445 trials
- Cost
- $247.76
- Cost / observed success
- $0.85
- Accuracy/cost frontier
- Dominated by another admitted system
- Top interval group
- Does not overlap leading interval
- Average duration
- 436.6 seconds per trial
- Tokens
- 47,653,156 input, 329,123,284 cached, 7,219,388 output
- Integrity review
- 2 disqualified trials, 0.45% adjustment
- Run date
- 2026-05-01
- Verified
- 2026-07-27
Environment: Harbor-managed Docker task containers using the official Terminal-Bench 2.1 dataset.
Network: Task-defined network access; passing trajectories were reviewed for solution lookup and reward hacking.
Budget: Official per-task time and resource limits with no timeout or resource overrides; five trials per task.
Admission fields
- Exact harness version and model
- Benchmark and dataset version
- Reasoning effort and token usage
- Sandbox, network, and budget policy
- Tasks, attempts, and total trials
- Cost, duration, run date, and primary source
- Integrity review and disqualified trials
Claim boundary
The same harness can move when the model, effort, version, task set, budget, or integrity policy changes.
HarnessMatch never copies this accuracy into workflow fit, operational readiness, or code auditability.
The displayed 95% interval is a normal approximation from the reported standard error. It is descriptive, not a task-cluster-corrected equivalence test.
Research behind the policy
These papers define reproducibility, episode traces, cost-aware evaluation, versioning, and contamination controls.