Coding Agent Performance Varies Greatly by Harness, Not Just Model

Naman Vats, Oleg Golev· July 28, 2026 View original

Summary

A study reveals that the performance and efficiency of AI coding agents are heavily influenced by the "harness" (the surrounding framework managing tools and context), not just the underlying language model. Different harnesses can lead to up to a 40x difference in token usage for solved tasks, even with minimal pass-rate variations.

Public leaderboards for AI coding agents often focus solely on the underlying language model's name and its pass rate, overlooking the significant impact of the "harness" – the framework that orchestrates tool usage, context management, and execution flow. This research demonstrates that when the harness varies, the observed performance and efficiency metrics conflate the capabilities of the model with the design of its surrounding scaffold. Experiments comparing models like Qwen 3.6 Plus and MiniMax M2.5 across different open-source harnesses showed dramatic differences in token consumption, sometimes up to 40 times more tokens per solved task, while pass rates remained relatively similar. This indicates that the harness design profoundly affects real-world costs, latency, and the need for human oversight. The study recommends evaluating coding agents as "harness-model pairs" and reporting comprehensive metrics beyond just pass rates, including token usage and latency, along with full harness specifications.

Why it matters

Professionals deploying or evaluating AI coding agents must consider the entire system (model + harness) to accurately assess cost, performance, and operational efficiency, rather than just the base model.

How to implement this in your domain

  1. 1When selecting or developing coding agents, evaluate the complete "harness-model pair" rather than just the language model in isolation.
  2. 2Benchmark agent performance using metrics beyond pass rate, including token usage, latency, and the number of "no-action" turns.
  3. 3Document and share full harness specifications alongside any performance comparisons to ensure reproducibility and accurate evaluation.
  4. 4Prioritize harnesses that optimize for efficiency and reduce unnecessary token generation, even if pass rates are similar across different setups.

Who benefits

Software DevelopmentAI EngineeringDevOpsIT Consulting

Key takeaways

  • Coding agent performance is heavily influenced by the surrounding harness, not just the LLM.
  • Harness choice can lead to massive differences in token usage and operational cost.
  • Evaluate coding agents as "harness-model pairs" for accurate assessment.
  • Comprehensive metrics like token usage and latency are crucial for real-world deployment.

Original post by Naman Vats, Oleg Golev

"arXiv:2607.22585v1 Announce Type: new Abstract: Public leaderboards for coding agents typically rank systems by model name and pass rate, while the surrounding harness (the scaffold that issues tools, manages context, and decides when to stop) is often under-specified. Model-to-m…"

View on X

Originally posted by Naman Vats, Oleg Golev on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses