Methodology 3.3 / 2026-08-02

How the catalog is built.

HarnessMatch organizes source-backed facts for inspection and comparison. It does not select a winner or combine the catalog into one universal score.

1Define

Category membership follows four documented harness criteria.

2Record

Every positive product claim links to a dated first-party source.

3Separate

Capabilities, popularity, code access, and benchmark results stay distinct.

4Expose

Unknowns, limitations, and verification dates remain visible.

Browse the technical methodology

1. Question and scope

Goal: make coding harnesses easier to inspect, filter, and compare without hiding editorial assumptions inside a personalized ranking. The catalog records what a product is, which mechanisms are documented, where it runs, which public signals are observable, and what remains unknown.

Model intelligence is never treated as harness capability. Product mechanisms, public activity, documentation breadth, source-code access, and benchmark configurations remain separate throughout the data model and interface.

Monthly editorial spotlight

The homepage may feature up to three active harnesses chosen by HarnessMatch to expose contrasting product approaches worth investigating. The selection is editorial, alphabetical, and independent of popularity, source count, capability levels, or benchmark results. It is not a ranking, an overall recommendation, or evidence from comparative HarnessMatch product trials.

The month identifies the publication period, not a source-verification date. The quality gate rejects a build once the edition no longer matches the current UTC month. Names, descriptions, limitations, and check dates come from the current catalog records, while each linked profile retains the supporting first-party evidence and its scope.

2. Catalog membership

A product is classified as a coding harness only when all four conditions below have current first-party evidence. The boundary is adapted from What makes a harness a harness. That source is a single-author conceptual preprint, so HarnessMatch uses it as a conservative catalog rule rather than an accepted standard or performance model.

Adaptive agent loop

The system repeatedly observes results and chooses the next action instead of following a fixed one-pass graph.

Repository tool execution

The system can use tools to inspect and change a repository or its execution environment.

Task-aware context management

The runtime assembles, updates, compacts, retrieves, or persists task-relevant context while work proceeds.

Model-independent runtime control

Permissions, budgets, interruption, policy, or stop controls operate outside the model's own text generation.

Each criterion is recorded as documented, contradicted, or unknown. Four documented criteria establish category membership only. Unknown evidence is not converted into a negative capability claim.

Neighboring layers remain visible but separate

Coding harness

Owns the adaptive loop, repository tools, context handling, and runtime controls needed to perform coding work.

External harness orchestrator

Coordinates independent harnesses or user-supervised sessions but does not establish its own coding loop.

Framework or runtime

Supplies building blocks or durable execution for constructing agents rather than a ready-to-use coding harness.

Adjacent tool

Covers gateways, pure editor assistance, evaluation harnesses, and other useful systems outside the coding-harness boundary.

Layer and product role are independent. A platform can qualify as a coding harness when it owns the loop; a control plane that only supervises external harnesses does not. Every layer remains available in the catalog and comparison view.

Repository-local instructions, specifications, plans, progress files, handoff artifacts, skills, hooks, and verification scripts can form a valuable harness configuration around an agent. They remain adjacent artifacts for membership purposes unless the named product itself also owns the adaptive loop, repository tool execution, active context management, and model-independent runtime control.

Claims and public activity

Capabilities are stored as source-linked claims rather than inferred from product names or model providers. Claim states preserve whether a mechanism is available by default, documented, optional, surface-specific, not documented, explicitly absent, or deprecated. “Not documented” remains uncertainty; “no built-in support” is used only when an admitted source says so explicitly.

Reusable-skill support records whether the harness itself documents loading task-specific instruction packages such as Agent Skills or SKILL.md. It does not catalog individual skills, plugins, hooks, commands, or MCP servers. Support does not establish portability, package quality, safety, or adoption; third-party skills remain executable or trusted content to review separately.

Active, dormant, and archived product states remain explicit. Dormant and archived records may stay available for research continuity, while active catalog views and measured rankings exclude them where the page states that scope.

The Usage page records source-native observations from OpenRouter, Homebrew, npm, filtered GitHub release assets, VS Code Marketplace, Open VSX, JetBrains Marketplace, and GitHub repositories. Tokens, requests, downloads, installs, stars, forks, and release cadence observe different populations and denominators. They are never added together, used as capability evidence, or treated as a quality score. Missing coverage means no admitted mapping, not zero adoption. Ranking bars use a visible linear scale anchored to the largest mapped value in the selected source; focused comparisons retain that same source-wide scale and global source rank. A 1 px origin marker distinguishes positive values too small to occupy one display pixel without changing their calculated bar width. A source-coverage contract requires every active harness to be either exactly mapped or explicitly retained as unmapped for each of the eight source views.

The stable-release tracker is factual and separate. It joins reviewed product-specific tag patterns to canonical repositories, excludes drafts, prereleases, and unrelated release trains, and publishes the latest tag, date, official URL, repository scope, observation date, and trailing 90-day count. Release frequency is maintenance context rather than quality or task-success evidence.

3. GUI workflow classification

The GUI catalog describes a different product layer. A harness-native GUI exposes its own coding harness, while a multi-harness workspace supervises independent CLIs or agent runtimes. Neither is universally better, and an external control plane does not inherit the capabilities of the harnesses it launches.

One agent + review

Stay close to one coding session and inspect changes in a dedicated visual surface. Required: visual review. Preferred: workspace isolation.

Parallel local tasks

Run several coding tasks on one machine without their branches or files colliding. Required: parallel sessions, workspace isolation. Preferred: visual review.

Remote control

Start or revisit coding work away from the machine where the repository is running. Required: remote execution or access. Preferred: shared team access.

Shared team workspace

Let teammates enter the same live workspace to watch, steer, preview, or review work. Required: remote execution or access, shared team access. Preferred: workspace isolation, visual review.

Every required and preferred claim documented yields Strong fit. Documented requirements with an unresolved preferred claim yield Good fit. An unresolved requirement yields Conditional fit. Inactive products or contradicted requirements are Not eligible on current evidence.

Products are alphabetical inside each band. No numeric value, source count, popularity signal, price, or license contributes to GUI fit.

GUI profiles group each source by one primary presentation topic in this fixed order: Product and workflow, Harness integrations, Sessions, isolation and review, Remote and collaboration, Public code and implementation. Each group shows one source in record order and keeps every remaining source available behind a disclosure. Topic placement organizes the ledger only; it does not transfer harness capability, rank sources, or change workflow fit.

4. Capability rubric

The editorial capability rubric is an inspectable coding aid, not a product score. Each axis is ordinal and independent. Levels are never summed into a universal grade or presented as measured performance.

SimplicityView five behavioral anchors
  1. Level 1Routine use requires extensive manual setup or several coordinated components.
  2. Level 2The main path is documented, but setup or routine operation still requires substantial configuration.
  3. Level 3Installation and the normal task loop are documented, with several choices or dependencies to manage.
  4. Level 4A short setup path and focused defaults cover normal use; advanced configuration is optional.
  5. Level 5A direct onboarding path and focused defaults support useful work with minimal required configuration.
FlexibilityView five behavioral anchors
  1. Level 1One prescribed provider and workflow, with little documented extension.
  2. Level 2A small set of documented provider, tool, or workflow choices.
  3. Level 3Several documented choices within a primary workflow.
  4. Level 4Multiple providers plus documented tool or workflow extension points.
  5. Level 5Broad provider choice and programmable extension across tools, agents, and operating modes.
Execution safetyView five behavioral anchors
  1. Level 1Host execution with no documented approval or isolation mechanism.
  2. Level 2Basic confirmations or restrictions, while broad host access remains the normal posture.
  3. Level 3Documented approvals or optional isolation for sensitive actions.
  4. Level 4Granular permissions plus documented isolation or policy controls.
  5. Level 5Enforced policy and isolation controls with explicit boundaries for execution and access.
AutonomyView five behavioral anchors
  1. Level 1Turn-by-turn assistance that depends on continuous user direction.
  2. Level 2Short multi-step work with frequent user intervention.
  3. Level 3Multi-step tasks with documented planning, continuation, or session recovery.
  4. Level 4Longer task loops that can use tools with limited intervention.
  5. Level 5Durable autonomous orchestration with delegation, recovery, and an explicit task lifecycle.
AutomationView five behavioral anchors
  1. Level 1Interactive use only, with no documented programmatic entry point.
  2. Level 2Callable or scriptable operation that still expects an attended session.
  3. Level 3A documented headless or non-interactive execution path.
  4. Level 4CI, scheduled, or API-driven execution with structured inputs or outputs.
  5. Level 5Durable automation with triggers, orchestration, monitoring, and recovery mechanisms.
Large-repository supportView five behavioral anchors
  1. Level 1A narrow working set selected manually, with no documented repository-wide mechanism.
  2. Level 2Manual context selection or basic repository search supports broader work.
  3. Level 3Documented repository search and cross-file context support.
  4. Level 4Managed indexing, context compression, or delegation supports broad repository work.
  5. Level 5Repository-scale context has documented refresh, partitioning, or durable lifecycle mechanisms.
Human controlView five behavioral anchors
  1. Level 1Few documented controls for intervention, review, or recovery.
  2. Level 2Coarse confirmation or post-change review with limited reversal support.
  3. Level 3Approvals or diff review are documented for common actions.
  4. Level 4Granular approvals are combined with checkpoints, rollback, or scoped permissions.
  5. Level 5Policy-level controls provide auditable actions and a reversible change workflow.

Publishing anchors makes editorial judgment inspectable. It does not establish inter-rater reliability, construct validity, or predictive performance.

5. Seven architecture layers

Architecture is described one layer at a time and never summed into a universal readiness grade.

Execution & isolation

Source-backed ordinal description of the most advanced documented mechanism on this layer.

Tooling & integrations

Source-backed ordinal description of the most advanced documented mechanism on this layer.

Context & state

Source-backed ordinal description of the most advanced documented mechanism on this layer.

Lifecycle & recovery

Source-backed ordinal description of the most advanced documented mechanism on this layer.

Observability

Source-backed ordinal description of the most advanced documented mechanism on this layer.

Verification

Source-backed ordinal description of the most advanced documented mechanism on this layer.

Governance & permissions

Source-backed ordinal description of the most advanced documented mechanism on this layer.

Recovery remains visible inside lifecycle; permissions are interpreted as governance; execution and tooling are explicit instead of inferred from one overall label.

6. Evidence states

Documented

A first-party document, repository, or announcement directly supports the claim and records a verification date.

Code-verifiable

The relevant client or full source is public and inspected at a pinned commit.

Independently measured

A complete external benchmark configuration passes the metadata admission policy.

Replicated

Two independent, compatible measurements reproduce the claim. No current harness receives this state by default.

States are claim-specific and need not form a simple ladder. Documentation volume is shown as coverage context only; it does not increase capability or scientific confidence.

Harness profiles group first-party sources in this fixed presentation order: Product and interfaces, Execution and control, Agents, state and recovery, Automation and extensions, Enterprise and operations, Releases and public code audit, Additional first-party sources. Each group shows one source in record order and keeps every remaining source available behind a disclosure. This layout does not rank, weight, or increase the confidence of any source.

7. AI-assisted, source-governed research

HarnessMatch uses language models to help discover, extract, structure, and cross-check information from first-party documentation, official repositories, release notes, and benchmark records admitted by the benchmark policy.

Model output is a research aid, not evidence. Every published product claim must remain traceable to an admitted underlying source and a verification date; that source record, not the model response, is authoritative.

Different models may be used independently to surface conflicting interpretations and reduce single-model blind spots. Not every claim is processed by every model, and model agreement does not establish accuracy, product capability, or scientific validity. Conflicts and unsupported claims are held for editorial review.

Discover

Find candidate products, documentation, repositories, releases, and benchmark records.

Extract

Turn source material into structured candidate claims, dates, versions, and limitations.

Cross-check

Use independent passes to flag disagreements, missing context, and claims needing closer inspection.

Publish

Admit only claims with an allowed source, direct traceability, editorial classification, and a verification date.

AI assistance improves research coverage and update speed. It does not replace source provenance, editorial judgment, inter-rater validation, or direct measurement.

8. Public-code artifacts

48 official repositories are inspected at exact commits for Security policy, CI workflow, Automated tests, Evaluation assets, Contributor documentation. The interface reports a transparent count out of five, not a weighted product score.

Presence does not establish adequacy, security, maintainability, or benchmark independence. Support-only repositories are shown but remain unranked.

9. Measured systems and uncertainty

No result is admitted without model, exact harness version, benchmark version, budget, sandbox or environment, attempts, date, cost, and primary source. A result belongs to model × harness × configuration × environment × budget.

The current archive contains 1 benchmark family: 5 Terminal-Bench 2.1 configurations, each with 89 tasks, five attempts, and 445 trials. This is exploratory evidence, not a general harness leaderboard. HarnessMatch shows a descriptive 95% interval calculated as accuracy ± 1.96 × the reported standard error, marks interval overlap with the leading configuration, and identifies the non-dominated accuracy and cost Pareto frontier.

The normal approximation does not model task clustering or benchmark sampling. Overlap is a visual uncertainty group, not a formal equivalence test. A task-cluster bootstrap or generalized mixed model remains preferable when raw trial data is available.

Inspect the benchmark archive

10. Classification is descriptive

Catalog layer

Whether the product owns a coding-agent loop, coordinates external harnesses, supplies a framework or runtime, or is an adjacent tool. Only the first layer qualifies as a coding harness.

Product role

Whether the product is a focused pair programmer, a coding agent, a broader agent, an extensible harness, or a platform that hosts agents.

Agent organization

Whether work stays in one agent loop, can be delegated to subagents, or is coordinated by a multi-agent runtime.

Interaction surfaces

Where the product is actually available: terminal, IDE, web or desktop, and automation.

Runtime and isolation

Whether execution is host-, sandbox-, or managed-first, followed by the documented isolation paths available to the user.

State and recovery

Whether the product is session-based or maintains persistent agent memory, and whether file rollback or restore is first-class.

Model access

Whether access is tied to one vendor, supports multiple providers or local models, or targets enterprise routing.

Controls and verification

Approval posture, permissions, hooks, isolation, recovery, and verification signals are compared separately from model capability.

Runtime posture and available isolation remain separate. Multi-agent organization, autonomy, and a larger feature surface are not assumed to be universally better.

Model portability

Model portability is a categorical posture derived from the documented provider style and local-model path. It remains independent from product capability:

Vendor-specific

The documented product path is tied to one vendor's model access.

Managed routing

An organization controls the documented provider or deployment routes.

Provider choice

The harness documents more than one provider, without a current local-model claim.

Provider + local

The harness documents multiple providers and a local or self-hosted model path.

11. Validation status

Implemented

Source-governed membership, explicit catalog layers, public ordinal anchors, dated claims, pinned repository commits, complete benchmark metadata, descriptive intervals, and Pareto status.

Protocol published

A fixed stratified sample, independent coding procedure, agreement statistics, uncertainty reporting, and held-out recoding rule are defined in the repository.

Not yet established

Inter-rater reliability, external content-validity review, criterion validity against controlled outcomes, and usability with real comparison tasks.

Open the inter-rater validation protocol9 products, 7 axes, 2 independent raters

Fixed sample: Claude Code, Codex, OpenCode, Aider, OpenHands, Cursor CLI, Qwen Code, mini-SWE-agent, Coder Agents.

The fixed sample spans open and closed products, terminal and IDE workflows, host-first, sandbox-first, and managed runtimes, plus pair-programmer, coding-agent, and platform roles.

  1. Freeze the source packet and product revision before coding begins.
  2. Give both raters the same public anchors and source packet, without access to the other rater's assignments.
  3. Record one ordinal level plus the exact supporting source for every harness-axis unit.
  4. Calculate agreement before reconciliation; disagreements are preserved in the public audit artifact.
  5. Revise an ambiguous anchor, then independently recode a held-out set before changing production ratings.
Unit
one harness × one capability axis
Primary statistic
Ordinal Krippendorff's alpha reported separately for each capability axis
Secondary statistic
Quadratic-weighted Cohen's kappa by axis
Uncertainty
Bootstrap 95% intervals by resampling harnesses; pilot intervals are expected to be wide

The pre-specified working threshold is α ≥ 0.80. This is a pre-specified working threshold, not a universal law. The coefficient, interval, raw agreement, disagreements, and sample composition must all be reported, and no pooled result may hide a weak axis. Even high agreement would establish reliability, not construct validity.

Content validity study

External developers, maintainers, DevOps practitioners, security engineers, and experienced terminal and IDE agent users. Rate whether each dimension and behavioral anchor is relevant, clear, and sufficient for choosing a coding harness. Product preference is not requested.

Item-level relevance and clarity distributions; Panel composition and disagreements; Dimensions or anchors retained, revised, added, or removed.
Comparison usability study

Prospective pilot: give developers realistic comparison tasks, record which catalog and evidence views they use, and test whether they can explain the relevant trade-offs before choosing what to evaluate themselves.

Accuracy when identifying documented capabilities, unknowns, and evidence limits; Time and interaction cost required to build a shortlist; Decision confidence without implying that the site selected a winner; Comprehension of the boundary between popularity, documentation, and measured performance.

Use the first pilot to estimate completion rates and task variance, then set the confirmatory sample from a preregistered precision or power target. Pilot results alone do not establish general usability.

12. Operational reference values

The Data page can order complete operational profiles using five equally weighted axes: Context 20%, Permissions 20%, Verification 20%, Observability 20%, Recovery 20%. If any axis is unknown, the profile remains unranked.

ContextBasic 40, Managed 82, Persistent 100, Unknown unranked
PermissionsHost 35, Approval 75, Policy 100, Unknown unranked
VerificationManual 35, Tool assisted 75, Workflow gated 100, Unknown unranked
ObservabilitySession 45, Logs 78, Traces 100, Unknown unranked
RecoveryManual 25, Session resume 60, Checkpoint 85, Managed recovery 100, Unknown unranked

These values describe documented operational posture for one analytical view. They are not probabilities, benchmark outcomes, or universal product-quality grades.

Scientific basis

52 methodological sources inform measurement, uncertainty, benchmark validity, harness architecture, and claim limits. Papers guide the method; current first-party records still control product claims.

Open all 52 research sources
Handbook on Constructing Composite Indicators: Methodology and User GuideRequires a theoretical framework, explicit missing-data treatment, normalization, weighting, aggregation, and uncertainty and sensitivity analysis for composite indicators.Limit: The handbook addresses composite indicators broadly and does not supply harness-specific constructs or valid product ratings.OECD / European Commission JRC 2008, guidanceSoftware Modeling and Measurement: The Goal/Question/Metric ParadigmStructures measurement from a concrete goal through decision questions to metrics, preventing collection of signals that do not support a user decision.Limit: GQM structures the measurement program but does not validate a particular rubric, value function, or benchmark.University of Maryland Technical Report 1992, guidanceISO/IEC/IEEE 15939:2017 Systems and Software Engineering - Measurement ProcessRequires information needs, defined measures, collection and analysis procedures, and evaluation of whether measurement results are valid.Limit: The standard defines a measurement process rather than product-specific scoring rules for AI coding harnesses.ISO/IEC/IEEE Standard, standardApplying Inter-Rater Reliability and Agreement in Collaborative Grounded Theory Studies in Software EngineeringSupports dual coding, iterative codebook improvement, and transparent inter-rater reliability reporting for editorial software classifications.Limit: Reliability measures coder consistency; they do not by themselves establish that a construct predicts real harness outcomes.Journal of Systems and Software 2023krippendorffsalpha: An R Package for Measuring Agreement Using Krippendorff's Alpha CoefficientSupports agreement analysis for ordinal ratings, missing assignments, and more than two coders while emphasizing that interpretation thresholds depend on context.Limit: The coefficient measures reliability rather than validity; HarnessMatch must also report uncertainty, raw disagreement, and the composition of the coded sample.The R Journal 2021Establishing Best Practices in Building Rigorous Agentic BenchmarksSeparates task validity from outcome validity and requires frozen environments, verified solvability, robust graders, open harnesses, contamination controls, baselines, and uncertainty.Limit: The checklist evaluates benchmark rigor, not current product mechanisms or suitability for an individual workflow.NeurIPS 2025 Datasets and BenchmarksJudging LLM-as-a-Judge with MT-Bench and Chatbot ArenaDocuments position, verbosity, self-enhancement, and reasoning biases in model-based judging and validates judge agreement against controlled and crowdsourced human preferences.Limit: The study evaluates chat assistants rather than code patches. It supports caution and human calibration for BuffBench-style judges, not a Codebuff capability or quality rating.NeurIPS 2023 Datasets and BenchmarksHolistic Evaluation of Language ModelsArgues for standardized, multi-scenario, multi-metric evaluation that exposes trade-offs rather than collapsing every outcome into a single score.Limit: HELM evaluates language models; HarnessMatch applies its reporting principles while keeping model and harness capability separate.Transactions on Machine Learning Research 2023General Agent EvaluationProposes a unified protocol and full-factorial agent, model, and environment evaluation, reinforcing that general-agent claims require cross-domain testing rather than one coding score.Limit: It does not evaluate goose and remains a preprint; HarnessMatch uses it as methodology evidence, not as product evidence or a product rating.arXiv 2026, preprintWildClawBench: A Benchmark for Real-World, Long-Horizon Agent EvaluationUses native CLI harnesses, reproducible containers, human-authored long-horizon tasks, side-effect auditing, and hybrid grading; it also demonstrates that changing the harness can materially change outcomes for a fixed model.Limit: It is a recent preprint with 60 tasks. HarnessMatch does not import its Hermes Agent or other product results until exact harness revisions, model settings, budgets, attempts, and task-level records satisfy the benchmark admission policy.arXiv 2026, preprintAgents Don't Just Agree, They Remember: Benchmarking Persistent Sycophancy in Stateful Personal AgentsShows why persistent memory needs a separate governance layer: incorrect user claims can cross a durable write boundary and later reappear after the original session has been cleared.Limit: The study evaluates a specific persistent-sycophancy protocol on Hermes Agent and OpenClaw across twelve models; it does not establish a general product-quality score or prove that every memory write is unsafe.arXiv 2026, preprintSleeper Channels and Provenance Gates: Persistent Prompt Injection in Always-on Autonomous AI AgentsModels cross-surface delayed execution through memory, skills, schedules, and files, motivating provenance, one-shot authorization, and separation between in-process guardrails and OS-level containment for persistent agents.Limit: The formal argument and companion prototype remain a preprint; the end-to-end attack is demonstrated on OpenClaw, while Hermes Agent is part of the threat model rather than the evaluated implementation target.arXiv 2026, preprintYet Even Less Is Even Better For Agentic, Reasoning, and Coding LLMsStudies training-trajectory selection across mini-SWE-agent, MSWE-agent, and other scaffolds, reinforcing that the agent framework and trajectory construction must be recorded when interpreting coding-model results.Limit: The experiments optimize model training and use specific framework configurations; they do not compare current harness products for end-user suitability, so HarnessMatch imports no mini-SWE-agent product score from the paper.arXiv 2026, preprintLAGUNA M.1 / XS.2 Technical ReportRecords the production pool harness, model sampling parameters, per-task sandbox resources, four repeated runs, benchmark patches, and reward-hacking review, illustrating how harness and environment details change the interpretation of agentic model results.Limit: This is a vendor technical report centered on Poolside models. It does not pin the pool client revision or publish complete task-level trajectories for product comparison, and it modifies benchmark images, so HarnessMatch imports no Poolside Agent CLI score from it.Poolside technical report 2026, preprintExpanding the AI Evaluation Toolbox with Statistical ModelsDistinguishes fixed-benchmark accuracy from generalized accuracy and motivates statistical models that decompose task and system variation with valid uncertainty.Limit: HarnessMatch currently lacks raw task-level trials for a generalized linear mixed model and therefore labels its intervals descriptive.NIST AI 800-3 2026, standardSWE-bench: Can Language Models Resolve Real-World GitHub Issues?Establishes repository-level issue resolution with real GitHub issues, complete codebases, executable environments, and fail-to-pass tests as a more realistic unit than isolated code generation.Limit: The original benchmark covers twelve Python repositories and early model systems; later results still require current task curation, exact harness configuration, repeated attempts, and contamination controls.ICLR 2024SWE-agent: Agent-Computer Interfaces Enable Automated Software EngineeringTreats the interface around the model as a first-class variable: prompts, commands, control flow, environment, and feedback format.Limit: It studies SWE-agent and benchmark tasks; it is not a current capability audit of commercial products.NeurIPS 2024OpenHands: An Open Platform for AI Software Developers as Generalist AgentsSeparates the agent loop from the execution runtime and documents sandbox, event stream, tools, and multi-agent coordination.Limit: Its architecture is a strong reference model, not proof that every product implements the same boundaries.ICLR 2025Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line InterfacesRequires realistic tasks, pinned environments, executable tests, repeated attempts, and explicit agent-harness configuration.Limit: A score belongs to a particular model, harness, environment, budget, and run policy; not to the harness in isolation.ICLR 2026Harness Engineering for Agentic AI Coding Tools: An Exploratory StudyDefines eight repository-level mechanisms: context files, settings, skills, subagents, commands, hooks, rules, and MCP.Limit: The product matrix is a February 2026 snapshot of five tools, so current capability labels still need live first-party verification.AIware 2026Harness-Bench: Measuring Harness Effects across Models in Realistic Agent WorkflowsFrames a harness through context, tools, state, constraints, permissions, tracing, and recovery, and compares model-harness pairings.Limit: It is a recent preprint and its aggregate results are not imported as permanent product ratings.arXiv 2026, preprintDon't Blame the Large Language Model: How Agent Harness Evolution Shapes Coding Agent QualityHolds the model fixed across 35 Qwen Code releases, showing why the exact harness version must be recorded alongside any result.Limit: It is a recent preprint focused on one harness's release history and 50 SWE-bench Verified tasks, not a universal product ranking.arXiv 2026, preprintAgentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent HarnessesMakes component, experience, and decision observability explicit, pairs every harness edit with a falsifiable prediction, and tests transfer across benchmarks and model families.Limit: It evaluates an automatically evolved research harness and remains a preprint; its gains do not validate current commercial products or justify copying its aggregate scores into HarnessMatch.arXiv 2026, preprintRethinking the Evaluation of Harness Evolution for AgentsRequires matched feedback and inference budgets, simple test-time-scaling baselines, and held-out tasks when claiming that harness optimization improves general capability.Limit: It studies automatic harness evolution on Terminal-Bench 2.1 with two model families and is too recent to serve as an accepted universal evaluation standard.arXiv 2026, preprintWhat makes a harness a harness: necessary and sufficient conditions for an agent harnessProvides an operational inclusion test and separates an agent harness from an SDK, framework, IDE plugin, orchestrator, and the evaluation harness used to score it.Limit: It is a single-author conceptual-analysis preprint; the proposed boundary is useful for catalog consistency but is not yet a consensus taxonomy or performance model.arXiv 2026, preprintFailure as a Process: An Anatomy of CLI Coding Agent TrajectoriesAnalyzes complete Terminal-Bench trajectories and motivates explicit validation, recovery, and intervention mechanisms rather than final-score-only evaluation.Limit: It studies OpenHands, MiniSWE, and Terminus2 trajectories; its findings guide comparison axes but do not verify other products' capabilities.arXiv 2026, preprintAgent Harness Engineering: A SurveySeparates execution, tooling, context, lifecycle, observability, verification, and governance as distinct harness layers.Limit: It is under review; HarnessMatch uses the taxonomy as an audit framework rather than treating it as an accepted standard.TMLR submission 2026, preprintLarge Language Model-Based Agents for Software Engineering: A SurveySeparates model reasoning from agent perception, tools, action, memory, planning, and human or multi-agent interaction across 106 studies.Limit: It is a broad survey rather than an operational audit, so its categories guide vocabulary but do not verify current product features.arXiv 2024, preprintLLM-Based Multi-Agent Systems for Software Engineering: Literature Review, Vision, and the Road AheadShows that multi-agent topology, specialization, coordination, and human involvement are workflow choices rather than a universal quality ladder.Limit: The review covers multi-agent research systems; it does not establish that delegated agents improve every coding task or every product.ACM TOSEM 2025Understanding Software Engineering Agents: A Study of Thought-Action-Result TrajectoriesMotivates trajectory-level observability and distinguishes context gathering, editing, feedback use, and validation patterns hidden by final pass rates.Limit: The study analyzes 120 trajectories from three research agents and should not be generalized into product-level success rates.ASE 2025Agentless: Demystifying LLM-based Software Engineering AgentsProvides a strong counterexample to autonomy-as-quality: a simple localization, repair, and validation pipeline can be competitive and easier to inspect.Limit: Its results are benchmark- and model-specific, and they do not imply that interactive or long-horizon workflows never benefit from richer harnesses.PACMSE / FSE 2025AutoCodeRover: Autonomous Program ImprovementDemonstrates a structured repair loop combining repository search, program analysis, patching, and test feedback rather than undifferentiated autonomy.Limit: It targets issue resolution in a research system; its design supports the verification axis but not current commercial product claims.ISSTA 2024SWE-rebench: An Automated Pipeline for Task Collection and Decontaminated Evaluation of Software Engineering AgentsSupports continuously refreshed tasks, explicit evaluation windows, and contamination-aware reporting instead of permanent static leaderboard claims.Limit: Fresh tasks reduce contamination risk but do not remove the need to pin model, harness, budget, environment, attempts, and run policy.ICLR 2026REAP: Automatic Curation of Coding Agent Benchmarks from Interactive Production UsageDerives executable tasks from real developer-agent sessions and audits task relevance, test alignment, and multi-run stability, showing how production evaluation can preserve fidelity without relying on one static public benchmark.Limit: REAP studies one production monorepo pipeline and uses automated, partly model-assisted curation; its solve rates cannot be transferred to unrelated repositories or treated as harness-only ratings.arXiv 2026, preprintContext as a Tool: Context Management for Long-Horizon SWE-AgentsTreats context lifecycle as an explicit harness mechanism with stable task state, condensed long-term memory, working memory, and proactive compression.Limit: It evaluates one proposed context-management approach and does not justify assigning high context ratings from context-window size alone.arXiv 2025, preprintWink: Recovering from Misbehaviors in Coding AgentsMotivates recovery as a separate dimension by analyzing specification drift, reasoning problems, and tool-call failures in production trajectories.Limit: The reported recovery system is evaluated in one production setting and is not evidence that unrelated harnesses provide comparable intervention mechanisms.arXiv 2026, preprintAI Harness Engineering: A Runtime Substrate for Foundation-Model Software AgentsDefines eleven runtime responsibilities and treats an auditable execution episode, rather than a single response, as the evaluation unit.Limit: The proposed H0-H3 ladder is a new conceptual framework validated on a controlled task, not an accepted cross-product scoring standard.arXiv 2026, preprintCode as Agent HarnessOrganizes harness design into interface, mechanisms, and multi-agent scaling, with verification, shared state, regression control, and human oversight.Limit: It is a broad survey and roadmap, so its mechanisms guide comparison fields but do not supply current product capability labels.arXiv 2026, preprintVeRO: An Evaluation Harness for Agents to Optimize AgentsRequires versioned agent snapshots, controlled budgets, reference procedures, and structured execution traces for reproducible agent evaluation.Limit: VeRO studies agent optimization tasks and does not make its optimizer results directly comparable to general coding-harness leaderboards.arXiv 2026, preprintHolistic Agent Leaderboard: The Missing Infrastructure for AI Agent EvaluationDemonstrates standardized, cost-aware evaluation across 21,730 rollouts and uses log inspection to surface behavior hidden by aggregate success rates.Limit: Its cross-domain infrastructure validates evaluation practice, not a permanent ordering of the coding products in HarnessMatch.ICLR 2026R2E: Turning any GitHub Repository into a Programming Agent EnvironmentShows why repository evaluation needs executable environments, realistic project state, program analysis, and feedback rather than static code generation alone.Limit: R2E evaluates constructed repository environments and does not audit the production safety, recovery, or governance of commercial harnesses.ICML 2024Vision2Web: A Hierarchical Benchmark for Visual Website Development with Agent VerificationUses hierarchical tasks, executable test cases, GUI-agent verification, and visual judges to evaluate end-to-end work beyond patch-only benchmarks.Limit: The benchmark is specialized for website development and cannot be generalized into an overall coding-harness performance score.ICML 2026Claw-SWE-Bench: A Benchmark for Evaluating OpenClaw-style Agent Harnesses on Coding TasksQuantifies large model and harness effects under controlled comparisons and treats cost as a first-class evaluation outcome.Limit: It is a recent benchmark preprint; its measured pairings are not converted into harness-only ratings or extrapolated to unrelated workflows.arXiv 2026, preprintCopilot Evaluation Harness: Building User Trust in LLMs and LM Agents for IDE EnvironmentsCombines static and execution-based metrics across documentation, testing, and bug-fixing tasks, and examines prompt sensitivity and agent behavior.Limit: The workshop submission evaluates selected models and tasks inside one harness; it is methodological evidence, not a current product ranking.ICLR BuildingTrust Workshop 2025, preprintAgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM AgentsSeparates task utility from adversarial security for tool-using agents and evaluates prompt injection through untrusted tool output with realistic tasks, explicit security cases, and adaptive attacks.Limit: AgentDojo focuses on web and productivity tools rather than repository coding, so it informs HarnessMatch's security model without proving any coding harness is secure or insecure.NeurIPS 2024 Datasets and BenchmarksDo Coding Agents Understand Least-Privilege Authorization?AuthBench uses 120 realistic terminal tasks with human-reviewed file permissions and executable utility and attack validators, showing why least privilege needs harness-level policy rather than model intuition alone.Limit: It is a recent preprint about model-generated file policies, not an audit of Crush permissions or evidence that any approval prompt or hook is an effective security boundary.arXiv 2026, preprintOvereager Coding Agents: Measuring Out-of-Scope Actions on Benign TasksDefines out-of-scope action as an authorization failure distinct from prompt injection and sandbox escape, and uses paired consent conditions, audited tool calls, repeated runs, and multiple model-harness combinations.Limit: The benchmark evaluates selected May 2026 configurations and remains a preprint; HarnessMatch does not convert its product rates into permanent security scores without complete versioned run records.arXiv 2026, preprintHow Coding Agents Fail Their Users: A Large-Scale Analysis of Developer-Agent Misalignment in 20,574 Real-World SessionsAnalyzes 20,574 sessions from 1,639 repositories across CLI and IDE workflows and motivates visible constraints, correction paths, and interface-specific evaluation of agent behavior.Limit: It is an observational preprint based on developer pushback as visible evidence of misalignment; it does not estimate hidden failures or provide comparative harness success rates.arXiv 2026, preprintSWE-chat: Coding Agent Interactions From Real Users in the WildStudies 6,000 public coding-agent sessions and distinguishes agent-dominant vibe coding from human-authored workflows, supporting workflow-specific comparison and evaluation with real interaction traces.Limit: It is a living observational dataset and recent preprint; public-session selection, attribution, and survival into commits do not establish causal product quality or safety rankings.arXiv 2026, preprintHarness engineering: leveraging Codex in an agent-first worldDescribes repository knowledge as a versioned system of record, short navigation-first AGENTS.md guidance, worktree-local application observability, mechanically enforced architecture, executable plans, and recurring maintenance as parts of an agent-legible engineering environment.Limit: This is a self-reported five-month internal Codex case study without a matched control or independent replication. Its throughput and time estimates cannot be generalized into product capability or comparative performance claims.OpenAI Engineering 2026, guidanceEffective harnesses for long-running agentsMotivates explicit initialization, feature decomposition, progress artifacts, Git history, end-to-end testing, and clean handoffs so fresh sessions can continue work across context windows instead of relying on compaction alone.Limit: This is an Anthropic internal engineering experiment around the Claude Agent SDK, not a controlled cross-product benchmark. It publishes design observations rather than a replicated success-rate estimate.Anthropic Engineering 2025, guidanceHarness design for long-running application developmentExplores planner, generator, and evaluator separation; negotiated completion contracts; browser-based verification; structured handoffs; context resets versus compaction; and removing harness components as model capability changes.Limit: The reported solo and full-harness comparison used one prompt with radically different duration and cost, plus qualitative manual assessment. It does not isolate causality, supply a product benchmark, or justify a universal multi-agent advantage.Anthropic Engineering 2026, guidance