HarnessMatch organizes source-backed facts for inspection and comparison. It does not select a winner or combine the catalog into one universal score.
1Define
Category membership follows four documented harness criteria.
2Record
Every positive product claim links to a dated first-party source.
3Separate
Capabilities, popularity, code access, and benchmark results stay distinct.
4Expose
Unknowns, limitations, and verification dates remain visible.
Browse the technical methodology
1. Question and scope
Goal: make coding harnesses easier to inspect, filter, and compare without hiding editorial assumptions inside a personalized ranking. The catalog records what a product is, which mechanisms are documented, where it runs, which public signals are observable, and what remains unknown.
Model intelligence is never treated as harness capability. Product mechanisms, public activity, documentation breadth, source-code access, and benchmark configurations remain separate throughout the data model and interface.
Monthly editorial spotlight
The homepage may feature up to three active harnesses chosen by HarnessMatch to expose contrasting product approaches worth investigating. The selection is editorial, alphabetical, and independent of popularity, source count, capability levels, or benchmark results. It is not a ranking, an overall recommendation, or evidence from comparative HarnessMatch product trials.
The month identifies the publication period, not a source-verification date. The quality gate rejects a build once the edition no longer matches the current UTC month. Names, descriptions, limitations, and check dates come from the current catalog records, while each linked profile retains the supporting first-party evidence and its scope.
2. Catalog membership
A product is classified as a coding harness only when all four conditions below have current first-party evidence. The boundary is adapted from What makes a harness a harness. That source is a single-author conceptual preprint, so HarnessMatch uses it as a conservative catalog rule rather than an accepted standard or performance model.
Adaptive agent loop
The system repeatedly observes results and chooses the next action instead of following a fixed one-pass graph.
Repository tool execution
The system can use tools to inspect and change a repository or its execution environment.
Task-aware context management
The runtime assembles, updates, compacts, retrieves, or persists task-relevant context while work proceeds.
Model-independent runtime control
Permissions, budgets, interruption, policy, or stop controls operate outside the model's own text generation.
Each criterion is recorded as documented, contradicted, or unknown. Four documented criteria establish category membership only. Unknown evidence is not converted into a negative capability claim.
Neighboring layers remain visible but separate
Coding harness
Owns the adaptive loop, repository tools, context handling, and runtime controls needed to perform coding work.
External harness orchestrator
Coordinates independent harnesses or user-supervised sessions but does not establish its own coding loop.
Framework or runtime
Supplies building blocks or durable execution for constructing agents rather than a ready-to-use coding harness.
Adjacent tool
Covers gateways, pure editor assistance, evaluation harnesses, and other useful systems outside the coding-harness boundary.
Layer and product role are independent. A platform can qualify as a coding harness when it owns the loop; a control plane that only supervises external harnesses does not. Every layer remains available in the catalog and comparison view.
Repository-local instructions, specifications, plans, progress files, handoff artifacts, skills, hooks, and verification scripts can form a valuable harness configuration around an agent. They remain adjacent artifacts for membership purposes unless the named product itself also owns the adaptive loop, repository tool execution, active context management, and model-independent runtime control.
Claims and public activity
Capabilities are stored as source-linked claims rather than inferred from product names or model providers. Claim states preserve whether a mechanism is available by default, documented, optional, surface-specific, not documented, explicitly absent, or deprecated. “Not documented” remains uncertainty; “no built-in support” is used only when an admitted source says so explicitly.
Reusable-skill support records whether the harness itself documents loading task-specific instruction packages such as Agent Skills or SKILL.md. It does not catalog individual skills, plugins, hooks, commands, or MCP servers. Support does not establish portability, package quality, safety, or adoption; third-party skills remain executable or trusted content to review separately.
Active, dormant, and archived product states remain explicit. Dormant and archived records may stay available for research continuity, while active catalog views and measured rankings exclude them where the page states that scope.
The Usage page records source-native observations from OpenRouter, Homebrew, npm, filtered GitHub release assets, VS Code Marketplace, Open VSX, JetBrains Marketplace, and GitHub repositories. Tokens, requests, downloads, installs, stars, forks, and release cadence observe different populations and denominators. They are never added together, used as capability evidence, or treated as a quality score. Missing coverage means no admitted mapping, not zero adoption. Ranking bars use a visible linear scale anchored to the largest mapped value in the selected source; focused comparisons retain that same source-wide scale and global source rank. A 1 px origin marker distinguishes positive values too small to occupy one display pixel without changing their calculated bar width. A source-coverage contract requires every active harness to be either exactly mapped or explicitly retained as unmapped for each of the eight source views.
The stable-release tracker is factual and separate. It joins reviewed product-specific tag patterns to canonical repositories, excludes drafts, prereleases, and unrelated release trains, and publishes the latest tag, date, official URL, repository scope, observation date, and trailing 90-day count. Release frequency is maintenance context rather than quality or task-success evidence.
3. GUI workflow classification
The GUI catalog describes a different product layer. A harness-native GUI exposes its own coding harness, while a multi-harness workspace supervises independent CLIs or agent runtimes. Neither is universally better, and an external control plane does not inherit the capabilities of the harnesses it launches.
One agent + review
Stay close to one coding session and inspect changes in a dedicated visual surface. Required: visual review. Preferred: workspace isolation.
Parallel local tasks
Run several coding tasks on one machine without their branches or files colliding. Required: parallel sessions, workspace isolation. Preferred: visual review.
Remote control
Start or revisit coding work away from the machine where the repository is running. Required: remote execution or access. Preferred: shared team access.
Shared team workspace
Let teammates enter the same live workspace to watch, steer, preview, or review work. Required: remote execution or access, shared team access. Preferred: workspace isolation, visual review.
Every required and preferred claim documented yields Strong fit. Documented requirements with an unresolved preferred claim yield Good fit. An unresolved requirement yields Conditional fit. Inactive products or contradicted requirements are Not eligible on current evidence.
Products are alphabetical inside each band. No numeric value, source count, popularity signal, price, or license contributes to GUI fit.
GUI profiles group each source by one primary presentation topic in this fixed order: Product and workflow, Harness integrations, Sessions, isolation and review, Remote and collaboration, Public code and implementation. Each group shows one source in record order and keeps every remaining source available behind a disclosure. Topic placement organizes the ledger only; it does not transfer harness capability, rank sources, or change workflow fit.
4. Capability rubric
The editorial capability rubric is an inspectable coding aid, not a product score. Each axis is ordinal and independent. Levels are never summed into a universal grade or presented as measured performance.
SimplicityView five behavioral anchors
Level 1Routine use requires extensive manual setup or several coordinated components.
Level 2The main path is documented, but setup or routine operation still requires substantial configuration.
Level 3Installation and the normal task loop are documented, with several choices or dependencies to manage.
Level 4A short setup path and focused defaults cover normal use; advanced configuration is optional.
Level 5A direct onboarding path and focused defaults support useful work with minimal required configuration.
FlexibilityView five behavioral anchors
Level 1One prescribed provider and workflow, with little documented extension.
Level 2A small set of documented provider, tool, or workflow choices.
Level 3Several documented choices within a primary workflow.
Level 4Multiple providers plus documented tool or workflow extension points.
Level 5Broad provider choice and programmable extension across tools, agents, and operating modes.
Execution safetyView five behavioral anchors
Level 1Host execution with no documented approval or isolation mechanism.
Level 2Basic confirmations or restrictions, while broad host access remains the normal posture.
Level 3Documented approvals or optional isolation for sensitive actions.
Level 4Granular permissions plus documented isolation or policy controls.
Level 5Enforced policy and isolation controls with explicit boundaries for execution and access.
AutonomyView five behavioral anchors
Level 1Turn-by-turn assistance that depends on continuous user direction.
Level 2Short multi-step work with frequent user intervention.
Level 3Multi-step tasks with documented planning, continuation, or session recovery.
Level 4Longer task loops that can use tools with limited intervention.
Level 5Durable autonomous orchestration with delegation, recovery, and an explicit task lifecycle.
AutomationView five behavioral anchors
Level 1Interactive use only, with no documented programmatic entry point.
Level 2Callable or scriptable operation that still expects an attended session.
Level 3A documented headless or non-interactive execution path.
Level 4CI, scheduled, or API-driven execution with structured inputs or outputs.
Level 5Durable automation with triggers, orchestration, monitoring, and recovery mechanisms.
Large-repository supportView five behavioral anchors
Level 1A narrow working set selected manually, with no documented repository-wide mechanism.
Level 5Repository-scale context has documented refresh, partitioning, or durable lifecycle mechanisms.
Human controlView five behavioral anchors
Level 1Few documented controls for intervention, review, or recovery.
Level 2Coarse confirmation or post-change review with limited reversal support.
Level 3Approvals or diff review are documented for common actions.
Level 4Granular approvals are combined with checkpoints, rollback, or scoped permissions.
Level 5Policy-level controls provide auditable actions and a reversible change workflow.
Publishing anchors makes editorial judgment inspectable. It does not establish inter-rater reliability, construct validity, or predictive performance.
5. Seven architecture layers
Architecture is described one layer at a time and never summed into a universal readiness grade.
Execution & isolation
Source-backed ordinal description of the most advanced documented mechanism on this layer.
Tooling & integrations
Source-backed ordinal description of the most advanced documented mechanism on this layer.
Context & state
Source-backed ordinal description of the most advanced documented mechanism on this layer.
Lifecycle & recovery
Source-backed ordinal description of the most advanced documented mechanism on this layer.
Observability
Source-backed ordinal description of the most advanced documented mechanism on this layer.
Verification
Source-backed ordinal description of the most advanced documented mechanism on this layer.
Governance & permissions
Source-backed ordinal description of the most advanced documented mechanism on this layer.
Recovery remains visible inside lifecycle; permissions are interpreted as governance; execution and tooling are explicit instead of inferred from one overall label.
6. Evidence states
Documented
A first-party document, repository, or announcement directly supports the claim and records a verification date.
Code-verifiable
The relevant client or full source is public and inspected at a pinned commit.
Independently measured
A complete external benchmark configuration passes the metadata admission policy.
Replicated
Two independent, compatible measurements reproduce the claim. No current harness receives this state by default.
States are claim-specific and need not form a simple ladder. Documentation volume is shown as coverage context only; it does not increase capability or scientific confidence.
Harness profiles group first-party sources in this fixed presentation order: Product and interfaces, Execution and control, Agents, state and recovery, Automation and extensions, Enterprise and operations, Releases and public code audit, Additional first-party sources. Each group shows one source in record order and keeps every remaining source available behind a disclosure. This layout does not rank, weight, or increase the confidence of any source.
7. AI-assisted, source-governed research
HarnessMatch uses language models to help discover, extract, structure, and cross-check information from first-party documentation, official repositories, release notes, and benchmark records admitted by the benchmark policy.
Model output is a research aid, not evidence. Every published product claim must remain traceable to an admitted underlying source and a verification date; that source record, not the model response, is authoritative.
Different models may be used independently to surface conflicting interpretations and reduce single-model blind spots. Not every claim is processed by every model, and model agreement does not establish accuracy, product capability, or scientific validity. Conflicts and unsupported claims are held for editorial review.
Discover
Find candidate products, documentation, repositories, releases, and benchmark records.
Extract
Turn source material into structured candidate claims, dates, versions, and limitations.
Cross-check
Use independent passes to flag disagreements, missing context, and claims needing closer inspection.
Publish
Admit only claims with an allowed source, direct traceability, editorial classification, and a verification date.
AI assistance improves research coverage and update speed. It does not replace source provenance, editorial judgment, inter-rater validation, or direct measurement.
8. Public-code artifacts
48 official repositories are inspected at exact commits for Security policy, CI workflow, Automated tests, Evaluation assets, Contributor documentation. The interface reports a transparent count out of five, not a weighted product score.
Presence does not establish adequacy, security, maintainability, or benchmark independence. Support-only repositories are shown but remain unranked.
9. Measured systems and uncertainty
No result is admitted without model, exact harness version, benchmark version, budget, sandbox or environment, attempts, date, cost, and primary source. A result belongs to model × harness × configuration × environment × budget.
The current archive contains 1 benchmark family: 5 Terminal-Bench 2.1 configurations, each with 89 tasks, five attempts, and 445 trials. This is exploratory evidence, not a general harness leaderboard. HarnessMatch shows a descriptive 95% interval calculated as accuracy ± 1.96 × the reported standard error, marks interval overlap with the leading configuration, and identifies the non-dominated accuracy and cost Pareto frontier.
The normal approximation does not model task clustering or benchmark sampling. Overlap is a visual uncertainty group, not a formal equivalence test. A task-cluster bootstrap or generalized mixed model remains preferable when raw trial data is available.
Whether the product owns a coding-agent loop, coordinates external harnesses, supplies a framework or runtime, or is an adjacent tool. Only the first layer qualifies as a coding harness.
Product role
Whether the product is a focused pair programmer, a coding agent, a broader agent, an extensible harness, or a platform that hosts agents.
Agent organization
Whether work stays in one agent loop, can be delegated to subagents, or is coordinated by a multi-agent runtime.
Interaction surfaces
Where the product is actually available: terminal, IDE, web or desktop, and automation.
Runtime and isolation
Whether execution is host-, sandbox-, or managed-first, followed by the documented isolation paths available to the user.
State and recovery
Whether the product is session-based or maintains persistent agent memory, and whether file rollback or restore is first-class.
Model access
Whether access is tied to one vendor, supports multiple providers or local models, or targets enterprise routing.
Controls and verification
Approval posture, permissions, hooks, isolation, recovery, and verification signals are compared separately from model capability.
Runtime posture and available isolation remain separate. Multi-agent organization, autonomy, and a larger feature surface are not assumed to be universally better.
Model portability
Model portability is a categorical posture derived from the documented provider style and local-model path. It remains independent from product capability:
Vendor-specific
The documented product path is tied to one vendor's model access.
Managed routing
An organization controls the documented provider or deployment routes.
Provider choice
The harness documents more than one provider, without a current local-model claim.
Provider + local
The harness documents multiple providers and a local or self-hosted model path.
A fixed stratified sample, independent coding procedure, agreement statistics, uncertainty reporting, and held-out recoding rule are defined in the repository.
Not yet established
Inter-rater reliability, external content-validity review, criterion validity against controlled outcomes, and usability with real comparison tasks.
Open the inter-rater validation protocol9 products, 7 axes, 2 independent raters
The fixed sample spans open and closed products, terminal and IDE workflows, host-first, sandbox-first, and managed runtimes, plus pair-programmer, coding-agent, and platform roles.
Freeze the source packet and product revision before coding begins.
Give both raters the same public anchors and source packet, without access to the other rater's assignments.
Record one ordinal level plus the exact supporting source for every harness-axis unit.
Calculate agreement before reconciliation; disagreements are preserved in the public audit artifact.
Revise an ambiguous anchor, then independently recode a held-out set before changing production ratings.
Unit
one harness × one capability axis
Primary statistic
Ordinal Krippendorff's alpha reported separately for each capability axis
Secondary statistic
Quadratic-weighted Cohen's kappa by axis
Uncertainty
Bootstrap 95% intervals by resampling harnesses; pilot intervals are expected to be wide
The pre-specified working threshold is α ≥ 0.80. This is a pre-specified working threshold, not a universal law. The coefficient, interval, raw agreement, disagreements, and sample composition must all be reported, and no pooled result may hide a weak axis. Even high agreement would establish reliability, not construct validity.
Content validity study
External developers, maintainers, DevOps practitioners, security engineers, and experienced terminal and IDE agent users. Rate whether each dimension and behavioral anchor is relevant, clear, and sufficient for choosing a coding harness. Product preference is not requested.
Item-level relevance and clarity distributions; Panel composition and disagreements; Dimensions or anchors retained, revised, added, or removed.
Comparison usability study
Prospective pilot: give developers realistic comparison tasks, record which catalog and evidence views they use, and test whether they can explain the relevant trade-offs before choosing what to evaluate themselves.
Accuracy when identifying documented capabilities, unknowns, and evidence limits; Time and interaction cost required to build a shortlist; Decision confidence without implying that the site selected a winner; Comprehension of the boundary between popularity, documentation, and measured performance.
Use the first pilot to estimate completion rates and task variance, then set the confirmatory sample from a preregistered precision or power target. Pilot results alone do not establish general usability.
12. Operational reference values
The Data page can order complete operational profiles using five equally weighted axes: Context 20%, Permissions 20%, Verification 20%, Observability 20%, Recovery 20%. If any axis is unknown, the profile remains unranked.
These values describe documented operational posture for one analytical view. They are not probabilities, benchmark outcomes, or universal product-quality grades.
Scientific basis
52 methodological sources inform measurement, uncertainty, benchmark validity, harness architecture, and claim limits. Papers guide the method; current first-party records still control product claims.