We first check whether a tool can meet your requirements, then compare workflow fit and show how stable the result is. Sources and measured configurations remain separate from editorial judgments.
1Scope
Only products that own a documented coding loop enter the default ranking.
2Requirements
Must-haves are gates. A strength elsewhere cannot cancel a missing requirement.
3Workflow fit
Eligible tools are compared against your preferences, not a universal quality score.
4Uncertainty
Sensitivity and evidence states show how much confidence to place in the order.
Browse the technical methodology
1. Decision question: what should this person use?
Goal: select a coding harness for a declared workflow. Questions: can it satisfy the non-negotiable constraints; among eligible tools, which mechanisms fit the workflow; how stable is that ordering; and what evidence supports each claim? Metrics: eligibility, preference value, rank robustness, evidence state, and measured system outcomes.
Model intelligence is never treated as harness capability. Product mechanisms and benchmark configurations remain separate throughout the data model and interface.
2. Scope and eligibility before preference
Catalog membership test
A product enters the default coding-harness recommender only when all four conditions below have current first-party evidence. The boundary is adapted from What makes a harness a harness. That source is a single-author conceptual preprint, so HarnessMatch uses it as a conservative catalog rule rather than an accepted standard or performance model.
Adaptive agent loop
The system repeatedly observes results and chooses the next action instead of following a fixed one-pass graph.
Repository tool execution
The system can use tools to inspect and change a repository or its execution environment.
Task-aware context management
The runtime assembles, updates, compacts, retrieves, or persists task-relevant context while work proceeds.
Model-independent runtime control
Permissions, budgets, interruption, policy, or stop controls operate outside the model's own text generation.
Each criterion is recorded as documented, contradicted, or unknown. Only four documented criteria pass. Unknown evidence is never converted into a negative capability claim, but it still blocks ranking because category membership has not been established.
Neighboring layers remain visible but separate
Coding harness
Owns the adaptive loop, repository tools, context handling, and runtime controls needed to perform coding work.
External harness orchestrator
Coordinates independent harnesses or user-supervised sessions but does not establish its own coding loop.
Framework or runtime
Supplies building blocks or durable execution for constructing agents rather than a ready-to-use coding harness.
Adjacent tool
Covers gateways, pure editor assistance, evaluation harnesses, and other useful systems outside the coding-harness boundary.
Layer and product role are independent. A platform can still qualify as a coding harness when it owns the loop; a control plane that only supervises external harnesses does not. Orchestrators, frameworks, and adjacent tools remain available to catalog and compare, but they do not enter the default recommendation ordering.
Workflow gates
After membership, interface, model-access path, explicit required features, and mode-implied requirements are non-compensatory gates. Consumer subscription access, enterprise access, provider breadth, and local-model support are recorded independently: one never establishes another by inference. Choosing no model-access preference skips that gate instead of inventing a default constraint. CI requires headless execution; parallel work requires documented subagents. A high value elsewhere cannot compensate for a failed gate.
Each capability is stored once as a source-linked claim; catalog filters and eligibility gates are derived from that record rather than maintained as separate yes/no fields. Claims preserve operating state: available by default, documented, optional, surface-specific, not documented, explicitly absent, or deprecated. A supported gate must link to a first-party source and verification date. “Not documented” remains uncertainty; “no built-in support” is used only when an admitted source says so explicitly.
Only active products are eligible: dormant and archived products remain visible for research but are excluded from recommendations and benchmark rankings. OpenRouter and GitHub are discovery sources only: they can create a research candidate, but cannot establish a capability.
The current public status is deliberately conservative: “Eligible” means catalog membership and every declared workflow gate have current supporting documentation. “Not eligible on current evidence” can mean a neighboring product layer or at least one undocumented gate; it does not prove technical impossibility.
3. Preference model and swing weights
Eligible products are compared with a provisional linear additive MCDA model. Interface and model access are absent from the score because they were already used as gates. The published reference swing weights sum to 100:
Top priority30%
Control style25%
Change scope25%
Operating mode20%
The editorial 1-5 rubric is mapped through the explicit provisional value function 1→0, 2→25, 3→50, 4→75, 5→100. These internal values enable ordering; the public interface reports preference bands and rank robustness instead of presenting them as measured “quality out of 100”.
Control style
Approval heavyHuman Control 55%, Permissions 30%, Recovery 15%
Editors assign the most advanced level supported by first-party records. These anchors define the coding rule; they are not benchmark outcomes, model assessments, or validated performance scales.
SimplicityView five behavioral anchors
Level 1Routine use requires extensive manual setup or several coordinated components.
Level 2The main path is documented, but setup or routine operation still requires substantial configuration.
Level 3Installation and the normal task loop are documented, with several choices or dependencies to manage.
Level 4A short setup path and focused defaults cover normal use; advanced configuration is optional.
Level 5A direct onboarding path and focused defaults support useful work with minimal required configuration.
FlexibilityView five behavioral anchors
Level 1One prescribed provider and workflow, with little documented extension.
Level 2A small set of documented provider, tool, or workflow choices.
Level 3Several documented choices within a primary workflow.
Level 4Multiple providers plus documented tool or workflow extension points.
Level 5Broad provider choice and programmable extension across tools, agents, and operating modes.
Execution safetyView five behavioral anchors
Level 1Host execution with no documented approval or isolation mechanism.
Level 2Basic confirmations or restrictions, while broad host access remains the normal posture.
Level 3Documented approvals or optional isolation for sensitive actions.
Level 4Granular permissions plus documented isolation or policy controls.
Level 5Enforced policy and isolation controls with explicit boundaries for execution and access.
AutonomyView five behavioral anchors
Level 1Turn-by-turn assistance that depends on continuous user direction.
Level 2Short multi-step work with frequent user intervention.
Level 3Multi-step tasks with documented planning, continuation, or session recovery.
Level 4Longer task loops that can use tools with limited intervention.
Level 5Durable autonomous orchestration with delegation, recovery, and an explicit task lifecycle.
AutomationView five behavioral anchors
Level 1Interactive use only, with no documented programmatic entry point.
Level 2Callable or scriptable operation that still expects an attended session.
Level 3A documented headless or non-interactive execution path.
Level 4CI, scheduled, or API-driven execution with structured inputs or outputs.
Level 5Durable automation with triggers, orchestration, monitoring, and recovery mechanisms.
Large-repository supportView five behavioral anchors
Level 1A narrow working set selected manually, with no documented repository-wide mechanism.
Level 5Repository-scale context has documented refresh, partitioning, or durable lifecycle mechanisms.
Human controlView five behavioral anchors
Level 1Few documented controls for intervention, review, or recovery.
Level 2Coarse confirmation or post-change review with limited reversal support.
Level 3Approvals or diff review are documented for common actions.
Level 4Granular approvals are combined with checkpoints, rollback, or scoped permissions.
Level 5Policy-level controls provide auditable actions and a reversible change workflow.
Publishing an anchor makes an editorial judgment inspectable, not automatically reliable. Claim-level source mapping, independent dual coding, and external validity checks remain open work.
The additive model assumes the displayed criteria are sufficiently preference-independent for this use. That assumption is provisional and is listed as a validation target below.
4. Sensitivity instead of false precision
For each answer set, HarnessMatch evaluates 512 deterministic sensitivity scenarios. Every reference weight is multiplied by a value between 0.7 and 1.3, then renormalized; each provisional factor value is also stressed by up to ±12.5 points, half of one rubric step. The interface reports top-rank frequency, top-three frequency, mean rank, and best-worst rank.
Any result within 2 internal preference points of the highest provisional value is presented as part of the leading group. This display rule does not change the calculation; it prevents a small editorial-value difference from being presented as a uniquely superior product.
A “top three in 90%” result means 90% of tested preference-and-rating stress scenarios place the harness in the top three. It is not a 90% chance of task success and is not a Bayesian posterior. The uncertainty ranges are methodological stress bounds, not empirically estimated rating-error distributions.
If any scored operational input is undocumented, the candidate remains unranked for that comparison. Missing values are not renormalized away and are never converted into evidence of absence.
5. Seven architecture layers
Architecture is described one layer at a time and never summed into a universal readiness grade.
Execution & isolation
Source-backed ordinal description of the most advanced documented mechanism on this layer.
Tooling & integrations
Source-backed ordinal description of the most advanced documented mechanism on this layer.
Context & state
Source-backed ordinal description of the most advanced documented mechanism on this layer.
Lifecycle & recovery
Source-backed ordinal description of the most advanced documented mechanism on this layer.
Observability
Source-backed ordinal description of the most advanced documented mechanism on this layer.
Verification
Source-backed ordinal description of the most advanced documented mechanism on this layer.
Governance & permissions
Source-backed ordinal description of the most advanced documented mechanism on this layer.
Recovery is kept visible inside lifecycle; permissions are interpreted as governance; execution and tooling are now explicit instead of being inferred from a single “readiness” number.
6. Evidence states, not source-count confidence
Documented
A first-party document, repository, or announcement directly supports the claim and records a verification date.
Code-verifiable
The relevant client or full source is public and inspected at a pinned commit.
Independently measured
A complete external benchmark configuration passes the metadata admission policy.
Replicated
Two independent, compatible measurements reproduce the claim. No current harness is awarded this state by default.
States are claim-specific and need not form a simple ladder. The legacy high/medium/limited coverage label only describes documentation breadth: high requires 4 sources across 2 kinds; it is not used as scientific confidence.
Compact result cards summarize evidence available for the product record. They do not upgrade every individual feature claim to the strongest product-level state; the source rows on each profile remain authoritative.
7. AI-assisted, source-governed research
HarnessMatch uses language models to help discover, extract, structure, and cross-check information from first-party documentation, official repositories, release notes, and benchmark records admitted by the benchmark policy.
Model output is a research aid, not evidence. Every published product claim must remain traceable to an admitted underlying source and a verification date; that source record, not the model response, is authoritative.
Different models may be used independently to surface conflicting interpretations and reduce single-model blind spots. Not every claim is processed by every model, and model agreement does not establish accuracy, product capability, or scientific validity. Conflicts and unsupported claims are held for editorial review.
Discover
Find candidate products, documentation, repositories, releases, and benchmark records.
Extract
Turn source material into structured candidate claims, dates, versions, and limitations.
Cross-check
Use independent passes to flag disagreements, missing context, and claims needing closer inspection.
Publish
Admit only claims with an allowed source, direct traceability, editorial classification, and a verification date.
AI assistance improves research coverage and update speed. It does not replace source provenance, editorial judgment, inter-rater validation, or direct measurement.
8. Public-code artifacts
36 official repositories are inspected at exact commits for Security policy, CI workflow, Automated tests, Evaluation assets, Contributor documentation. The interface reports a transparent count out of five, not a weighted product score.
Presence does not establish adequacy, security, maintainability, or benchmark independence. Support-only repositories are shown but remain unranked.
9. Measured systems and uncertainty
No result is admitted without model, exact harness version, benchmark version, budget, sandbox/environment, attempts, date, cost, and primary source. A result belongs to model × harness × configuration × environment × budget.
The current view contains 5 Terminal-Bench 2.1 configurations, each with 89 tasks, five attempts, and 445 trials. HarnessMatch shows a descriptive 95% interval calculated as accuracy ± 1.96 × the reported standard error, marks interval overlap with the leading configuration, and identifies the non-dominated accuracy/cost Pareto frontier.
The normal approximation does not model task clustering or benchmark sampling. Overlap is a visual uncertainty group, not a formal equivalence test. A task-cluster bootstrap or generalized mixed model remains the preferred future analysis when raw trial data is available.
10. Classification is descriptive
Catalog layer
Whether the product owns a coding-agent loop, coordinates external harnesses, supplies a framework or runtime, or is an adjacent tool. Only the first layer enters the default recommender.
Product role
Whether the product is a focused pair programmer, a coding agent, a broader agent, an extensible harness, or a platform that hosts agents.
Agent organization
Whether work stays in one agent loop, can be delegated to subagents, or is coordinated by a multi-agent runtime.
Interaction surfaces
Where the product is actually available: terminal, IDE, web or desktop, and automation.
Runtime and isolation
Whether execution is host-, sandbox-, or managed-first, followed by the documented isolation paths available to the user.
State and recovery
Whether the product is session-based or maintains persistent agent memory, and whether file rollback or restore is first-class.
Model access
Whether access is tied to one vendor, supports multiple providers or local models, or targets enterprise routing.
Controls and verification
Approval posture, permissions, hooks, isolation, recovery, and verification signals are compared separately from model capability.
Runtime posture and available isolation remain separate. Multi-agent organization, autonomy, and a larger feature surface are not assumed to be universally better.
Workflow fit × model portability
The recommender result adds a two-dimensional reading aid, not another score. Rows reuse the user-specific workflow-fit bands. Columns derive a categorical model-portability posture from the separately documented provider style and local-model path:
Vendor-specific
The documented product path is tied to one vendor's model access.
Managed routing
An organization controls the documented provider or deployment routes.
Provider choice
The harness documents more than one provider, without a current local-model claim.
Provider + local
The harness documents multiple providers and a local or self-hosted model path.
Column placement does not add points, change the ordering, or claim that provider freedom is universally preferable. It lets a user see the trade-off between the fit calculated from their answers and the model-access posture they are willing to accept.
A fixed stratified sample, independent coding procedure, agreement statistics, uncertainty reporting, and held-out recoding rule are now defined in the repository.
Not yet established
Inter-rater reliability, content-validity review by external experts, criterion validity against controlled harness outcomes, and predictive validity in real user adoption.
Open the inter-rater validation protocol9 products, 7 axes, 2 independent raters
The fixed sample spans open and closed products, terminal and IDE workflows, host-first, sandbox-first, and managed runtimes, plus pair-programmer, coding-agent, and platform roles.
Freeze the source packet and product revision before coding begins.
Give both raters the same public anchors and source packet, without access to the other rater's assignments.
Record one ordinal level plus the exact supporting source for every harness-axis unit.
Calculate agreement before reconciliation; disagreements are preserved in the public audit artifact.
Revise an ambiguous anchor, then independently recode a held-out set before changing production ratings.
Unit
one harness × one capability axis
Primary statistic
Ordinal Krippendorff's alpha reported separately for each capability axis
Secondary statistic
Quadratic-weighted Cohen's kappa by axis
Uncertainty
Bootstrap 95% intervals by resampling harnesses; pilot intervals are expected to be wide
The pre-specified working threshold is α ≥ 0.80. This is a pre-specified working threshold, not a universal law. The coefficient, interval, raw agreement, disagreements, and sample composition must all be reported, and no pooled result may hide a weak axis. Even high agreement would establish reliability, not construct validity.
Content validity study
External developers, maintainers, DevOps practitioners, security engineers, and experienced terminal and IDE agent users. Rate whether each dimension and behavioral anchor is relevant, clear, and sufficient for choosing a coding harness. Product preference is not requested.
Item-level relevance and clarity distributions; Panel composition and disagreements; Dimensions or anchors retained, revised, added, or removed.
Prospective user study
Prospective pilot: collect workflow preferences, lock the recommendation, counterbalance hands-on trials of two or three eligible harnesses, and keep the predicted winner hidden until pre-reveal ratings are complete.
Pre-reveal match between the locked recommendation and the user's preferred harness; Perceived recommendation quality and transparency; Choice confidence and satisfaction after real use; Switching frequency and reasons after a follow-up period.
Use the first pilot to estimate variance and completion rates, then set the confirmatory sample from a preregistered precision or power target. Pilot results alone do not establish predictive validity.
These are transparent provisional value functions used inside workflow preference calculations. They are not benchmark outcomes, probabilities, or universal quality grades.
Scientific basis
52 methodological sources define measurement, MCDA, uncertainty, benchmark validity, harness architecture, and claim limits. Papers guide the method; current first-party records still control product claims.