Methodology 2.5 / 2026-07-28

How recommendations are built.

We first check whether a tool can meet your requirements, then compare workflow fit and show how stable the result is. Sources and measured configurations remain separate from editorial judgments.

1Scope

Only products that own a documented coding loop enter the default ranking.

2Requirements

Must-haves are gates. A strength elsewhere cannot cancel a missing requirement.

3Workflow fit

Eligible tools are compared against your preferences, not a universal quality score.

4Uncertainty

Sensitivity and evidence states show how much confidence to place in the order.

Browse the technical methodology

1. Decision question: what should this person use?

Goal: select a coding harness for a declared workflow. Questions: can it satisfy the non-negotiable constraints; among eligible tools, which mechanisms fit the workflow; how stable is that ordering; and what evidence supports each claim? Metrics: eligibility, preference value, rank robustness, evidence state, and measured system outcomes.

Model intelligence is never treated as harness capability. Product mechanisms and benchmark configurations remain separate throughout the data model and interface.

2. Scope and eligibility before preference

Catalog membership test

A product enters the default coding-harness recommender only when all four conditions below have current first-party evidence. The boundary is adapted from What makes a harness a harness. That source is a single-author conceptual preprint, so HarnessMatch uses it as a conservative catalog rule rather than an accepted standard or performance model.

Adaptive agent loop

The system repeatedly observes results and chooses the next action instead of following a fixed one-pass graph.

Repository tool execution

The system can use tools to inspect and change a repository or its execution environment.

Task-aware context management

The runtime assembles, updates, compacts, retrieves, or persists task-relevant context while work proceeds.

Model-independent runtime control

Permissions, budgets, interruption, policy, or stop controls operate outside the model's own text generation.

Each criterion is recorded as documented, contradicted, or unknown. Only four documented criteria pass. Unknown evidence is never converted into a negative capability claim, but it still blocks ranking because category membership has not been established.

Neighboring layers remain visible but separate

Coding harness

Owns the adaptive loop, repository tools, context handling, and runtime controls needed to perform coding work.

External harness orchestrator

Coordinates independent harnesses or user-supervised sessions but does not establish its own coding loop.

Framework or runtime

Supplies building blocks or durable execution for constructing agents rather than a ready-to-use coding harness.

Adjacent tool

Covers gateways, pure editor assistance, evaluation harnesses, and other useful systems outside the coding-harness boundary.

Layer and product role are independent. A platform can still qualify as a coding harness when it owns the loop; a control plane that only supervises external harnesses does not. Orchestrators, frameworks, and adjacent tools remain available to catalog and compare, but they do not enter the default recommendation ordering.

Workflow gates

After membership, interface, model-access path, explicit required features, and mode-implied requirements are non-compensatory gates. Consumer subscription access, enterprise access, provider breadth, and local-model support are recorded independently: one never establishes another by inference. Choosing no model-access preference skips that gate instead of inventing a default constraint. CI requires headless execution; parallel work requires documented subagents. A high value elsewhere cannot compensate for a failed gate.

Each capability is stored once as a source-linked claim; catalog filters and eligibility gates are derived from that record rather than maintained as separate yes/no fields. Claims preserve operating state: available by default, documented, optional, surface-specific, not documented, explicitly absent, or deprecated. A supported gate must link to a first-party source and verification date. “Not documented” remains uncertainty; “no built-in support” is used only when an admitted source says so explicitly.

Only active products are eligible: dormant and archived products remain visible for research but are excluded from recommendations and benchmark rankings. OpenRouter and GitHub are discovery sources only: they can create a research candidate, but cannot establish a capability.

The current public status is deliberately conservative: “Eligible” means catalog membership and every declared workflow gate have current supporting documentation. “Not eligible on current evidence” can mean a neighboring product layer or at least one undocumented gate; it does not prove technical impossibility.

3. Preference model and swing weights

Eligible products are compared with a provisional linear additive MCDA model. Interface and model access are absent from the score because they were already used as gates. The published reference swing weights sum to 100:

Top priority30%
Control style25%
Change scope25%
Operating mode20%

The editorial 1-5 rubric is mapped through the explicit provisional value function 1→0, 2→25, 3→50, 4→75, 5→100. These internal values enable ordering; the public interface reports preference bands and rank robustness instead of presenting them as measured “quality out of 100”.

Control style

Approval heavyHuman Control 55%, Permissions 30%, Recovery 15%
BalancedHuman Control 30%, Autonomy 25%, Permissions 20%, Verification 15%, Recovery 10%
Hands offAutonomy 25%, Automation 20%, Verification 20%, Recovery 15%, Security 10%, Observability 10%

Change scope

FocusedSimplicity 70%, Human Control 30%
Cross fileFlexibility 60%, Context 40%
Large repoLarge Repo 55%, Context 45%

Operating mode

InteractiveSimplicity 40%, Human Control 35%, Recovery 25%
CiAutomation 35%, Verification 25%, Security 20%, Observability 20%
ParallelAutonomy 30%, Context 25%, Observability 25%, Recovery 20%

Provisional capability anchors

Editors assign the most advanced level supported by first-party records. These anchors define the coding rule; they are not benchmark outcomes, model assessments, or validated performance scales.

SimplicityView five behavioral anchors
  1. Level 1Routine use requires extensive manual setup or several coordinated components.
  2. Level 2The main path is documented, but setup or routine operation still requires substantial configuration.
  3. Level 3Installation and the normal task loop are documented, with several choices or dependencies to manage.
  4. Level 4A short setup path and focused defaults cover normal use; advanced configuration is optional.
  5. Level 5A direct onboarding path and focused defaults support useful work with minimal required configuration.
FlexibilityView five behavioral anchors
  1. Level 1One prescribed provider and workflow, with little documented extension.
  2. Level 2A small set of documented provider, tool, or workflow choices.
  3. Level 3Several documented choices within a primary workflow.
  4. Level 4Multiple providers plus documented tool or workflow extension points.
  5. Level 5Broad provider choice and programmable extension across tools, agents, and operating modes.
Execution safetyView five behavioral anchors
  1. Level 1Host execution with no documented approval or isolation mechanism.
  2. Level 2Basic confirmations or restrictions, while broad host access remains the normal posture.
  3. Level 3Documented approvals or optional isolation for sensitive actions.
  4. Level 4Granular permissions plus documented isolation or policy controls.
  5. Level 5Enforced policy and isolation controls with explicit boundaries for execution and access.
AutonomyView five behavioral anchors
  1. Level 1Turn-by-turn assistance that depends on continuous user direction.
  2. Level 2Short multi-step work with frequent user intervention.
  3. Level 3Multi-step tasks with documented planning, continuation, or session recovery.
  4. Level 4Longer task loops that can use tools with limited intervention.
  5. Level 5Durable autonomous orchestration with delegation, recovery, and an explicit task lifecycle.
AutomationView five behavioral anchors
  1. Level 1Interactive use only, with no documented programmatic entry point.
  2. Level 2Callable or scriptable operation that still expects an attended session.
  3. Level 3A documented headless or non-interactive execution path.
  4. Level 4CI, scheduled, or API-driven execution with structured inputs or outputs.
  5. Level 5Durable automation with triggers, orchestration, monitoring, and recovery mechanisms.
Large-repository supportView five behavioral anchors
  1. Level 1A narrow working set selected manually, with no documented repository-wide mechanism.
  2. Level 2Manual context selection or basic repository search supports broader work.
  3. Level 3Documented repository search and cross-file context support.
  4. Level 4Managed indexing, context compression, or delegation supports broad repository work.
  5. Level 5Repository-scale context has documented refresh, partitioning, or durable lifecycle mechanisms.
Human controlView five behavioral anchors
  1. Level 1Few documented controls for intervention, review, or recovery.
  2. Level 2Coarse confirmation or post-change review with limited reversal support.
  3. Level 3Approvals or diff review are documented for common actions.
  4. Level 4Granular approvals are combined with checkpoints, rollback, or scoped permissions.
  5. Level 5Policy-level controls provide auditable actions and a reversible change workflow.

Publishing an anchor makes an editorial judgment inspectable, not automatically reliable. Claim-level source mapping, independent dual coding, and external validity checks remain open work.

The additive model assumes the displayed criteria are sufficiently preference-independent for this use. That assumption is provisional and is listed as a validation target below.

4. Sensitivity instead of false precision

For each answer set, HarnessMatch evaluates 512 deterministic sensitivity scenarios. Every reference weight is multiplied by a value between 0.7 and 1.3, then renormalized; each provisional factor value is also stressed by up to ±12.5 points, half of one rubric step. The interface reports top-rank frequency, top-three frequency, mean rank, and best-worst rank.

Any result within 2 internal preference points of the highest provisional value is presented as part of the leading group. This display rule does not change the calculation; it prevents a small editorial-value difference from being presented as a uniquely superior product.

A “top three in 90%” result means 90% of tested preference-and-rating stress scenarios place the harness in the top three. It is not a 90% chance of task success and is not a Bayesian posterior. The uncertainty ranges are methodological stress bounds, not empirically estimated rating-error distributions.

If any scored operational input is undocumented, the candidate remains unranked for that comparison. Missing values are not renormalized away and are never converted into evidence of absence.

5. Seven architecture layers

Architecture is described one layer at a time and never summed into a universal readiness grade.

Execution & isolation

Source-backed ordinal description of the most advanced documented mechanism on this layer.

Tooling & integrations

Source-backed ordinal description of the most advanced documented mechanism on this layer.

Context & state

Source-backed ordinal description of the most advanced documented mechanism on this layer.

Lifecycle & recovery

Source-backed ordinal description of the most advanced documented mechanism on this layer.

Observability

Source-backed ordinal description of the most advanced documented mechanism on this layer.

Verification

Source-backed ordinal description of the most advanced documented mechanism on this layer.

Governance & permissions

Source-backed ordinal description of the most advanced documented mechanism on this layer.

Recovery is kept visible inside lifecycle; permissions are interpreted as governance; execution and tooling are now explicit instead of being inferred from a single “readiness” number.

6. Evidence states, not source-count confidence

Documented

A first-party document, repository, or announcement directly supports the claim and records a verification date.

Code-verifiable

The relevant client or full source is public and inspected at a pinned commit.

Independently measured

A complete external benchmark configuration passes the metadata admission policy.

Replicated

Two independent, compatible measurements reproduce the claim. No current harness is awarded this state by default.

States are claim-specific and need not form a simple ladder. The legacy high/medium/limited coverage label only describes documentation breadth: high requires 4 sources across 2 kinds; it is not used as scientific confidence.

Compact result cards summarize evidence available for the product record. They do not upgrade every individual feature claim to the strongest product-level state; the source rows on each profile remain authoritative.

7. AI-assisted, source-governed research

HarnessMatch uses language models to help discover, extract, structure, and cross-check information from first-party documentation, official repositories, release notes, and benchmark records admitted by the benchmark policy.

Model output is a research aid, not evidence. Every published product claim must remain traceable to an admitted underlying source and a verification date; that source record, not the model response, is authoritative.

Different models may be used independently to surface conflicting interpretations and reduce single-model blind spots. Not every claim is processed by every model, and model agreement does not establish accuracy, product capability, or scientific validity. Conflicts and unsupported claims are held for editorial review.

Discover

Find candidate products, documentation, repositories, releases, and benchmark records.

Extract

Turn source material into structured candidate claims, dates, versions, and limitations.

Cross-check

Use independent passes to flag disagreements, missing context, and claims needing closer inspection.

Publish

Admit only claims with an allowed source, direct traceability, editorial classification, and a verification date.

AI assistance improves research coverage and update speed. It does not replace source provenance, editorial judgment, inter-rater validation, or direct measurement.

8. Public-code artifacts

36 official repositories are inspected at exact commits for Security policy, CI workflow, Automated tests, Evaluation assets, Contributor documentation. The interface reports a transparent count out of five, not a weighted product score.

Presence does not establish adequacy, security, maintainability, or benchmark independence. Support-only repositories are shown but remain unranked.

9. Measured systems and uncertainty

No result is admitted without model, exact harness version, benchmark version, budget, sandbox/environment, attempts, date, cost, and primary source. A result belongs to model × harness × configuration × environment × budget.

The current view contains 5 Terminal-Bench 2.1 configurations, each with 89 tasks, five attempts, and 445 trials. HarnessMatch shows a descriptive 95% interval calculated as accuracy ± 1.96 × the reported standard error, marks interval overlap with the leading configuration, and identifies the non-dominated accuracy/cost Pareto frontier.

The normal approximation does not model task clustering or benchmark sampling. Overlap is a visual uncertainty group, not a formal equivalence test. A task-cluster bootstrap or generalized mixed model remains the preferred future analysis when raw trial data is available.

10. Classification is descriptive

Catalog layer

Whether the product owns a coding-agent loop, coordinates external harnesses, supplies a framework or runtime, or is an adjacent tool. Only the first layer enters the default recommender.

Product role

Whether the product is a focused pair programmer, a coding agent, a broader agent, an extensible harness, or a platform that hosts agents.

Agent organization

Whether work stays in one agent loop, can be delegated to subagents, or is coordinated by a multi-agent runtime.

Interaction surfaces

Where the product is actually available: terminal, IDE, web or desktop, and automation.

Runtime and isolation

Whether execution is host-, sandbox-, or managed-first, followed by the documented isolation paths available to the user.

State and recovery

Whether the product is session-based or maintains persistent agent memory, and whether file rollback or restore is first-class.

Model access

Whether access is tied to one vendor, supports multiple providers or local models, or targets enterprise routing.

Controls and verification

Approval posture, permissions, hooks, isolation, recovery, and verification signals are compared separately from model capability.

Runtime posture and available isolation remain separate. Multi-agent organization, autonomy, and a larger feature surface are not assumed to be universally better.

Workflow fit × model portability

The recommender result adds a two-dimensional reading aid, not another score. Rows reuse the user-specific workflow-fit bands. Columns derive a categorical model-portability posture from the separately documented provider style and local-model path:

Vendor-specific

The documented product path is tied to one vendor's model access.

Managed routing

An organization controls the documented provider or deployment routes.

Provider choice

The harness documents more than one provider, without a current local-model claim.

Provider + local

The harness documents multiple providers and a local or self-hosted model path.

Column placement does not add points, change the ordering, or claim that provider freedom is universally preferable. It lets a user see the trade-off between the fit calculated from their answers and the model-access posture they are willing to accept.

11. Validation status

Implemented

Source-governed membership gates, catalog layers, workflow gates, explicit weights, public provisional capability anchors, deterministic sensitivity analysis, missing-data non-renormalization, source dates, pinned repository commits, complete benchmark metadata, descriptive intervals, and Pareto status.

Protocol published

A fixed stratified sample, independent coding procedure, agreement statistics, uncertainty reporting, and held-out recoding rule are now defined in the repository.

Not yet established

Inter-rater reliability, content-validity review by external experts, criterion validity against controlled harness outcomes, and predictive validity in real user adoption.

Open the inter-rater validation protocol9 products, 7 axes, 2 independent raters

Fixed sample: Claude Code, Codex, OpenCode, Aider, OpenHands, Cursor CLI, Qwen Code, mini-SWE-agent, Coder Agents.

The fixed sample spans open and closed products, terminal and IDE workflows, host-first, sandbox-first, and managed runtimes, plus pair-programmer, coding-agent, and platform roles.

  1. Freeze the source packet and product revision before coding begins.
  2. Give both raters the same public anchors and source packet, without access to the other rater's assignments.
  3. Record one ordinal level plus the exact supporting source for every harness-axis unit.
  4. Calculate agreement before reconciliation; disagreements are preserved in the public audit artifact.
  5. Revise an ambiguous anchor, then independently recode a held-out set before changing production ratings.
Unit
one harness × one capability axis
Primary statistic
Ordinal Krippendorff's alpha reported separately for each capability axis
Secondary statistic
Quadratic-weighted Cohen's kappa by axis
Uncertainty
Bootstrap 95% intervals by resampling harnesses; pilot intervals are expected to be wide

The pre-specified working threshold is α ≥ 0.80. This is a pre-specified working threshold, not a universal law. The coefficient, interval, raw agreement, disagreements, and sample composition must all be reported, and no pooled result may hide a weak axis. Even high agreement would establish reliability, not construct validity.

Content validity study

External developers, maintainers, DevOps practitioners, security engineers, and experienced terminal and IDE agent users. Rate whether each dimension and behavioral anchor is relevant, clear, and sufficient for choosing a coding harness. Product preference is not requested.

Item-level relevance and clarity distributions; Panel composition and disagreements; Dimensions or anchors retained, revised, added, or removed.
Prospective user study

Prospective pilot: collect workflow preferences, lock the recommendation, counterbalance hands-on trials of two or three eligible harnesses, and keep the predicted winner hidden until pre-reveal ratings are complete.

Pre-reveal match between the locked recommendation and the user's preferred harness; Perceived recommendation quality and transparency; Choice confidence and satisfaction after real use; Switching frequency and reasons after a follow-up period.

Use the first pilot to estimate variance and completion rates, then set the confirmatory sample from a preregistered precision or power target. Pilot results alone do not establish predictive validity.

Published internal value tables

ContextBasic 40, Managed 82, Persistent 100, Unknown unranked
PermissionsHost 35, Approval 75, Policy 100, Unknown unranked
VerificationManual 35, Tool assisted 75, Workflow gated 100, Unknown unranked
ObservabilitySession 45, Logs 78, Traces 100, Unknown unranked
RecoveryManual 25, Session resume 60, Checkpoint 85, Managed recovery 100, Unknown unranked

These are transparent provisional value functions used inside workflow preference calculations. They are not benchmark outcomes, probabilities, or universal quality grades.

Scientific basis

52 methodological sources define measurement, MCDA, uncertainty, benchmark validity, harness architecture, and claim limits. Papers guide the method; current first-party records still control product claims.

Open all 52 research sources
An Introductory Guide to Multi-Criteria Decision AnalysisDefines hard minima before compensatory scoring, cardinal value functions, swing weighting, preference-independence assumptions, cost separation, and sensitivity analysis.Limit: It is practitioner guidance for general decisions, so HarnessMatch must still validate criteria and preferences with coding-harness users.UK Government Analysis Function 2024, guidanceHandbook on Constructing Composite Indicators: Methodology and User GuideRequires a theoretical framework, explicit missing-data treatment, normalization, weighting, aggregation, and uncertainty and sensitivity analysis for composite indicators.Limit: The handbook addresses composite indicators broadly and does not supply harness-specific constructs or valid product ratings.OECD / European Commission JRC 2008, guidanceSoftware Modeling and Measurement: The Goal/Question/Metric ParadigmStructures measurement from a concrete goal through decision questions to metrics, preventing collection of signals that do not support a user decision.Limit: GQM structures the measurement program but does not validate a particular rubric, value function, or benchmark.University of Maryland Technical Report 1992, guidanceISO/IEC/IEEE 15939:2017 Systems and Software Engineering - Measurement ProcessRequires information needs, defined measures, collection and analysis procedures, and evaluation of whether measurement results are valid.Limit: The standard defines a measurement process rather than product-specific scoring rules for AI coding harnesses.ISO/IEC/IEEE Standard, standardStochastic Multicriteria Acceptability Analysis for Group Decision MakingMotivates rank-acceptability analysis when weights or performance inputs are uncertain or missing instead of publishing one brittle ordering.Limit: HarnessMatch currently varies weights only and should not call its deterministic sensitivity sample a full SMAA implementation.European Journal of Operational Research 2007Applying Inter-Rater Reliability and Agreement in Collaborative Grounded Theory Studies in Software EngineeringSupports dual coding, iterative codebook improvement, and transparent inter-rater reliability reporting for editorial software classifications.Limit: Reliability measures coder consistency; they do not by themselves establish that a construct predicts real harness outcomes.Journal of Systems and Software 2023krippendorffsalpha: An R Package for Measuring Agreement Using Krippendorff's Alpha CoefficientSupports agreement analysis for ordinal ratings, missing assignments, and more than two coders while emphasizing that interpretation thresholds depend on context.Limit: The coefficient measures reliability rather than validity; HarnessMatch must also report uncertainty, raw disagreement, and the composition of the coded sample.The R Journal 2021A User-Centric Evaluation Framework for Recommender SystemsMotivates evaluating perceived recommendation quality, transparency, usefulness, satisfaction, and behavioral intention instead of relying only on offline ranking behavior.Limit: ResQue was developed for recommender-system user experience, not coding harness selection; HarnessMatch still needs a domain-specific prospective study.ACM RecSys 2011Establishing Best Practices in Building Rigorous Agentic BenchmarksSeparates task validity from outcome validity and requires frozen environments, verified solvability, robust graders, open harnesses, contamination controls, baselines, and uncertainty.Limit: The checklist evaluates benchmark rigor, not current product mechanisms or suitability for an individual workflow.NeurIPS 2025 Datasets and BenchmarksJudging LLM-as-a-Judge with MT-Bench and Chatbot ArenaDocuments position, verbosity, self-enhancement, and reasoning biases in model-based judging and validates judge agreement against controlled and crowdsourced human preferences.Limit: The study evaluates chat assistants rather than code patches. It supports caution and human calibration for BuffBench-style judges, not a Codebuff capability or quality rating.NeurIPS 2023 Datasets and BenchmarksHolistic Evaluation of Language ModelsArgues for standardized, multi-scenario, multi-metric evaluation that exposes trade-offs rather than collapsing every outcome into a single score.Limit: HELM evaluates language models; HarnessMatch applies its reporting principles while keeping model and harness capability separate.Transactions on Machine Learning Research 2023General Agent EvaluationProposes a unified protocol and full-factorial agent, model, and environment evaluation, reinforcing that general-agent claims require cross-domain testing rather than one coding score.Limit: It does not evaluate goose and remains a preprint; HarnessMatch uses it as methodology evidence, not as product evidence or a product rating.arXiv 2026, preprintWildClawBench: A Benchmark for Real-World, Long-Horizon Agent EvaluationUses native CLI harnesses, reproducible containers, human-authored long-horizon tasks, side-effect auditing, and hybrid grading; it also demonstrates that changing the harness can materially change outcomes for a fixed model.Limit: It is a recent preprint with 60 tasks. HarnessMatch does not import its Hermes Agent or other product results until exact harness revisions, model settings, budgets, attempts, and task-level records satisfy the benchmark admission policy.arXiv 2026, preprintAgents Don't Just Agree, They Remember: Benchmarking Persistent Sycophancy in Stateful Personal AgentsShows why persistent memory needs a separate governance layer: incorrect user claims can cross a durable write boundary and later reappear after the original session has been cleared.Limit: The study evaluates a specific persistent-sycophancy protocol on Hermes Agent and OpenClaw across twelve models; it does not establish a general product-quality score or prove that every memory write is unsafe.arXiv 2026, preprintSleeper Channels and Provenance Gates: Persistent Prompt Injection in Always-on Autonomous AI AgentsModels cross-surface delayed execution through memory, skills, schedules, and files, motivating provenance, one-shot authorization, and separation between in-process guardrails and OS-level containment for persistent agents.Limit: The formal argument and companion prototype remain a preprint; the end-to-end attack is demonstrated on OpenClaw, while Hermes Agent is part of the threat model rather than the evaluated implementation target.arXiv 2026, preprintYet Even Less Is Even Better For Agentic, Reasoning, and Coding LLMsStudies training-trajectory selection across mini-SWE-agent, MSWE-agent, and other scaffolds, reinforcing that the agent framework and trajectory construction must be recorded when interpreting coding-model results.Limit: The experiments optimize model training and use specific framework configurations; they do not compare current harness products for end-user workflow fit, so HarnessMatch imports no mini-SWE-agent product score from the paper.arXiv 2026, preprintLAGUNA M.1 / XS.2 Technical ReportRecords the production pool harness, model sampling parameters, per-task sandbox resources, four repeated runs, benchmark patches, and reward-hacking review, illustrating how harness and environment details change the interpretation of agentic model results.Limit: This is a vendor technical report centered on Poolside models. It does not pin the pool client revision or publish complete task-level trajectories for product comparison, and it modifies benchmark images, so HarnessMatch imports no Poolside Agent CLI score from it.Poolside technical report 2026, preprintExpanding the AI Evaluation Toolbox with Statistical ModelsDistinguishes fixed-benchmark accuracy from generalized accuracy and motivates statistical models that decompose task and system variation with valid uncertainty.Limit: HarnessMatch currently lacks raw task-level trials for a generalized linear mixed model and therefore labels its intervals descriptive.NIST AI 800-3 2026, standardSWE-bench: Can Language Models Resolve Real-World GitHub Issues?Establishes repository-level issue resolution with real GitHub issues, complete codebases, executable environments, and fail-to-pass tests as a more realistic unit than isolated code generation.Limit: The original benchmark covers twelve Python repositories and early model systems; later results still require current task curation, exact harness configuration, repeated attempts, and contamination controls.ICLR 2024SWE-agent: Agent-Computer Interfaces Enable Automated Software EngineeringTreats the interface around the model as a first-class variable: prompts, commands, control flow, environment, and feedback format.Limit: It studies SWE-agent and benchmark tasks; it is not a current capability audit of commercial products.NeurIPS 2024OpenHands: An Open Platform for AI Software Developers as Generalist AgentsSeparates the agent loop from the execution runtime and documents sandbox, event stream, tools, and multi-agent coordination.Limit: Its architecture is a strong reference model, not proof that every product implements the same boundaries.ICLR 2025Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line InterfacesRequires realistic tasks, pinned environments, executable tests, repeated attempts, and explicit agent-harness configuration.Limit: A score belongs to a particular model, harness, environment, budget, and run policy; not to the harness in isolation.ICLR 2026Harness Engineering for Agentic AI Coding Tools: An Exploratory StudyDefines eight repository-level mechanisms: context files, settings, skills, subagents, commands, hooks, rules, and MCP.Limit: The product matrix is a February 2026 snapshot of five tools, so current capability labels still need live first-party verification.AIware 2026Harness-Bench: Measuring Harness Effects across Models in Realistic Agent WorkflowsFrames a harness through context, tools, state, constraints, permissions, tracing, and recovery, and compares model-harness pairings.Limit: It is a recent preprint and its aggregate results are not imported as permanent product ratings.arXiv 2026, preprintDon't Blame the Large Language Model: How Agent Harness Evolution Shapes Coding Agent QualityHolds the model fixed across 35 Qwen Code releases, showing why the exact harness version must be recorded alongside any result.Limit: It is a recent preprint focused on one harness's release history and 50 SWE-bench Verified tasks, not a universal product ranking.arXiv 2026, preprintAgentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent HarnessesMakes component, experience, and decision observability explicit, pairs every harness edit with a falsifiable prediction, and tests transfer across benchmarks and model families.Limit: It evaluates an automatically evolved research harness and remains a preprint; its gains do not validate current commercial products or justify copying its aggregate scores into HarnessMatch.arXiv 2026, preprintRethinking the Evaluation of Harness Evolution for AgentsRequires matched feedback and inference budgets, simple test-time-scaling baselines, and held-out tasks when claiming that harness optimization improves general capability.Limit: It studies automatic harness evolution on Terminal-Bench 2.1 with two model families and is too recent to serve as an accepted universal evaluation standard.arXiv 2026, preprintWhat makes a harness a harness: necessary and sufficient conditions for an agent harnessProvides an operational inclusion test and separates an agent harness from an SDK, framework, IDE plugin, orchestrator, and the evaluation harness used to score it.Limit: It is a single-author conceptual-analysis preprint; the proposed boundary is useful for catalog consistency but is not yet a consensus taxonomy or performance model.arXiv 2026, preprintFailure as a Process: An Anatomy of CLI Coding Agent TrajectoriesAnalyzes complete Terminal-Bench trajectories and motivates explicit validation, recovery, and intervention mechanisms rather than final-score-only evaluation.Limit: It studies OpenHands, MiniSWE, and Terminus2 trajectories; its findings guide comparison axes but do not verify other products' capabilities.arXiv 2026, preprintAgent Harness Engineering: A SurveySeparates execution, tooling, context, lifecycle, observability, verification, and governance as distinct harness layers.Limit: It is under review; HarnessMatch uses the taxonomy as an audit framework rather than treating it as an accepted standard.TMLR submission 2026, preprintLarge Language Model-Based Agents for Software Engineering: A SurveySeparates model reasoning from agent perception, tools, action, memory, planning, and human or multi-agent interaction across 106 studies.Limit: It is a broad survey rather than an operational audit, so its categories guide vocabulary but do not verify current product features.arXiv 2024, preprintLLM-Based Multi-Agent Systems for Software Engineering: Literature Review, Vision, and the Road AheadShows that multi-agent topology, specialization, coordination, and human involvement are workflow choices rather than a universal quality ladder.Limit: The review covers multi-agent research systems; it does not establish that delegated agents improve every coding task or every product.ACM TOSEM 2025Understanding Software Engineering Agents: A Study of Thought-Action-Result TrajectoriesMotivates trajectory-level observability and distinguishes context gathering, editing, feedback use, and validation patterns hidden by final pass rates.Limit: The study analyzes 120 trajectories from three research agents and should not be generalized into product-level success rates.ASE 2025Agentless: Demystifying LLM-based Software Engineering AgentsProvides a strong counterexample to autonomy-as-quality: a simple localization, repair, and validation pipeline can be competitive and easier to inspect.Limit: Its results are benchmark- and model-specific, and they do not imply that interactive or long-horizon workflows never benefit from richer harnesses.PACMSE / FSE 2025AutoCodeRover: Autonomous Program ImprovementDemonstrates a structured repair loop combining repository search, program analysis, patching, and test feedback rather than undifferentiated autonomy.Limit: It targets issue resolution in a research system; its design supports the verification axis but not current commercial product claims.ISSTA 2024SWE-rebench: An Automated Pipeline for Task Collection and Decontaminated Evaluation of Software Engineering AgentsSupports continuously refreshed tasks, explicit evaluation windows, and contamination-aware reporting instead of permanent static leaderboard claims.Limit: Fresh tasks reduce contamination risk but do not remove the need to pin model, harness, budget, environment, attempts, and run policy.ICLR 2026REAP: Automatic Curation of Coding Agent Benchmarks from Interactive Production UsageDerives executable tasks from real developer-agent sessions and audits task relevance, test alignment, and multi-run stability, showing how production evaluation can preserve fidelity without relying on one static public benchmark.Limit: REAP studies one production monorepo pipeline and uses automated, partly model-assisted curation; its solve rates cannot be transferred to unrelated repositories or treated as harness-only ratings.arXiv 2026, preprintContext as a Tool: Context Management for Long-Horizon SWE-AgentsTreats context lifecycle as an explicit harness mechanism with stable task state, condensed long-term memory, working memory, and proactive compression.Limit: It evaluates one proposed context-management approach and does not justify assigning high context ratings from context-window size alone.arXiv 2025, preprintWink: Recovering from Misbehaviors in Coding AgentsMotivates recovery as a separate dimension by analyzing specification drift, reasoning problems, and tool-call failures in production trajectories.Limit: The reported recovery system is evaluated in one production setting and is not evidence that unrelated harnesses provide comparable intervention mechanisms.arXiv 2026, preprintAI Harness Engineering: A Runtime Substrate for Foundation-Model Software AgentsDefines eleven runtime responsibilities and treats an auditable execution episode, rather than a single response, as the evaluation unit.Limit: The proposed H0-H3 ladder is a new conceptual framework validated on a controlled task, not an accepted cross-product scoring standard.arXiv 2026, preprintCode as Agent HarnessOrganizes harness design into interface, mechanisms, and multi-agent scaling, with verification, shared state, regression control, and human oversight.Limit: It is a broad survey and roadmap, so its mechanisms guide comparison fields but do not supply current product capability labels.arXiv 2026, preprintVeRO: An Evaluation Harness for Agents to Optimize AgentsRequires versioned agent snapshots, controlled budgets, reference procedures, and structured execution traces for reproducible agent evaluation.Limit: VeRO studies agent optimization tasks and does not make its optimizer results directly comparable to general coding-harness leaderboards.arXiv 2026, preprintHolistic Agent Leaderboard: The Missing Infrastructure for AI Agent EvaluationDemonstrates standardized, cost-aware evaluation across 21,730 rollouts and uses log inspection to surface behavior hidden by aggregate success rates.Limit: Its cross-domain infrastructure validates evaluation practice, not a permanent ordering of the coding products in HarnessMatch.ICLR 2026R2E: Turning any GitHub Repository into a Programming Agent EnvironmentShows why repository evaluation needs executable environments, realistic project state, program analysis, and feedback rather than static code generation alone.Limit: R2E evaluates constructed repository environments and does not audit the production safety, recovery, or governance of commercial harnesses.ICML 2024Vision2Web: A Hierarchical Benchmark for Visual Website Development with Agent VerificationUses hierarchical tasks, executable test cases, GUI-agent verification, and visual judges to evaluate end-to-end work beyond patch-only benchmarks.Limit: The benchmark is specialized for website development and cannot be generalized into an overall coding-harness performance score.ICML 2026Claw-SWE-Bench: A Benchmark for Evaluating OpenClaw-style Agent Harnesses on Coding TasksQuantifies large model and harness effects under controlled comparisons and treats cost as a first-class evaluation outcome.Limit: It is a recent benchmark preprint; its measured pairings are not converted into harness-only ratings or extrapolated to unrelated workflows.arXiv 2026, preprintCopilot Evaluation Harness: Building User Trust in LLMs and LM Agents for IDE EnvironmentsCombines static and execution-based metrics across documentation, testing, and bug-fixing tasks, and examines prompt sensitivity and agent behavior.Limit: The workshop submission evaluates selected models and tasks inside one harness; it is methodological evidence, not a current product ranking.ICLR BuildingTrust Workshop 2025, preprintAgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM AgentsSeparates task utility from adversarial security for tool-using agents and evaluates prompt injection through untrusted tool output with realistic tasks, explicit security cases, and adaptive attacks.Limit: AgentDojo focuses on web and productivity tools rather than repository coding, so it informs HarnessMatch's security model without proving any coding harness is secure or insecure.NeurIPS 2024 Datasets and BenchmarksDo Coding Agents Understand Least-Privilege Authorization?AuthBench uses 120 realistic terminal tasks with human-reviewed file permissions and executable utility and attack validators, showing why least privilege needs harness-level policy rather than model intuition alone.Limit: It is a recent preprint about model-generated file policies, not an audit of Crush permissions or evidence that any approval prompt or hook is an effective security boundary.arXiv 2026, preprintOvereager Coding Agents: Measuring Out-of-Scope Actions on Benign TasksDefines out-of-scope action as an authorization failure distinct from prompt injection and sandbox escape, and uses paired consent conditions, audited tool calls, repeated runs, and multiple model-harness combinations.Limit: The benchmark evaluates selected May 2026 configurations and remains a preprint; HarnessMatch does not convert its product rates into permanent security scores without complete versioned run records.arXiv 2026, preprintHow Coding Agents Fail Their Users: A Large-Scale Analysis of Developer-Agent Misalignment in 20,574 Real-World SessionsAnalyzes 20,574 sessions from 1,639 repositories across CLI and IDE workflows and motivates visible constraints, correction paths, and interface-specific evaluation of agent behavior.Limit: It is an observational preprint based on developer pushback as visible evidence of misalignment; it does not estimate hidden failures or provide comparative harness success rates.arXiv 2026, preprintSWE-chat: Coding Agent Interactions From Real Users in the WildStudies 6,000 public coding-agent sessions and distinguishes agent-dominant vibe coding from human-authored workflows, supporting workflow-specific recommendations and evaluation with real interaction traces.Limit: It is a living observational dataset and recent preprint; public-session selection, attribution, and survival into commits do not establish causal product quality or safety rankings.arXiv 2026, preprint