ArXiv Research: Quantifying the Harness Effect in Agentic Coding Systems
By Mr.Xu
Published: · 10 views
Summary:ArXiv has released a study quantifying the impact of harnesses (tools, prompts, and control flows) on the performance of agentic coding systems. The research compares Claude-agent SDK and DeepAgents on Claude Opus 4.8 and GPT-5.5 across 80 tasks, revealing that harnesses do not confer a significant average advantage. Claude Opus 4.8 showed a performance deficit on repository tasks but excelled in contest tasks, while GPT-5.5 maintained a stable performance. The study also provides a detailed cos
Background and Objective
An agentic coding system integrates a language model with a harness (tools, prompts, and control flows) to transform a chat model into an autonomous software engineer. Vendors typically tune harnesses for their own models, and practitioners often assume that this native pairing enhances task-solving capabilities. This study aims to quantify the impact of harnesses on model performance and explore how different task types influence harness effectiveness.
Methodology
The research utilizes a private, contamination-controlled dataset of 256 repository and contest tasks to compare the following models and harnesses:
- Claude-agent SDK with Claude Opus 4.8
- OpenAI Codex SDK with GPT-5.5
- Gemini-3.5-Flash and DeepSeek-V3.2 as control groups
A total of 800 runs were planned, with 792 graded by an isolated oracle.
Key Findings
-
Harness Impact on Model Performance is Not Significant:
- Claude Opus 4.8 showed a performance deficit of 9.0 percentage points on repository tasks but led by 23.7 percentage points on contest tasks (label-permutation p = 0.003).
- GPT-5.5 maintained a stable performance, with the harness having an insignificant impact (+1.25 percentage points, 95% CI [-4.4, +6.9]).
-
Influence of Task Types:
- The disparity in performance between repository and contest tasks indicates that task types significantly affect harness effectiveness.
-
Cost Analysis:
- On Claude Opus 4.8, the neutral harness cost 1.3 to 1.6 times the original usage cost per solved task.
- On GPT-5.5, this ratio was 1.2 times.
Conclusion and Recommendations
The study concludes that harnesses do not significantly enhance model performance, and developers should select harnesses based on task types. Additionally, the research suggests that more nuanced cost-benefit analyses are needed when evaluating harness effectiveness.
Industry Impact
This study provides AI developers with valuable insights into the relationship between harnesses and model performance, aiding in the optimization of AI system design and deployment. The findings also highlight the potential for further innovation in AI harness development to better integrate harnesses with models and improve overall AI system performance.
— END —Source: ArXiv AI (cs.AI) (2026-09-14)
Tags: #Agentic Coding #AI Research #Model Evaluation #Harness #Performance Analysis
Community Comments