Nvidia research: the harness matters more than the model

LLMsAI Agents
Illustration generated by AI for this article

The Core · TL;DR

  • Nvidia's custom harness pushed Claude Opus 5 from a 30% score to 100% on the ARC-AGI-3 benchmark without changing the model itself.
  • The harness combined a memory management system with a 'supervisor' component to guide long-horizon reasoning.
  • OpenAI separately found that tweaking two harness settings tripled its models' scores on the same benchmark, though without reaching 100%.
  • Microsoft's April research found all 19 LLMs it tested made significant errors on long-horizon document-editing tasks, regardless of model.

Claude Opus 5 scored 30% on ARC-AGI-3 running onone, the best result any model managed unassisted. Wrapped in a custom harness built by Nvidia researchers, that same model hit 100%.

Nothing about the underlying weights changed between those two runs. What changed was the scaffolding around it: a memory management system and a "supervisor" component that structured how the model approached the benchmark's long-horizon reasoning tasks.

Nvidia published the findings on Friday, and the conclusion is blunt. For tasks that unfold over many steps, the architecture wrapped around a model can matter more than the model itself. Adel El Hallack, VP of Product in Nvidia's AI unit, frames an AI agent as four parts working together: the model, the harness (or scaffolding), the runtime, and the skills and libraries the agent can call on. Nvidia's results suggest that of those four, the harness is doing outsized work.

The pattern isn't unique to Nvidia's setup. OpenAI ran its own tests last month and found that adjusting just two settings in its harness tripled its models' scores on the same ARC-AGI-3 benchmark. OpenAI's tuned models still fell short of the 100% Nvidia reached, but the direction of the result lines up: small changes to the surrounding system produced outsized jumps in performance, independent of any model upgrade.

Why long-horizon tasks expose the gap

ARC-AGI-3 is built to test interactive, multi-step reasoning rather than single-shot question answering, which is where harness design tends to matter most. Models have to track state, recover from mistakes, and plan across many turns, and a well-built supervisor layer can catch errors and redirect the model before they compound.

Microsoft's research from April backs up why this matters beyond one benchmark. Testing 19 LLMs on long-horizon document-editing tasks, Microsoft found that every model tested, frontier systems included, produced documents riddled with errors. The failure showed up regardless of which model was doing the editing.

Taken together, these results push back against the idea that benchmark leadership is purely a function of model quality. If a harness can more than triple a score without touching the weights, then comparing raw model performance in isolation understates what real-world deployments can achieve, and overstates how much of the gap is actually about the model.

For teams building production AI agents, the implication is practical rather than academic: investment in scaffolding, memory systems, and supervisory logic may deliver bigger returns than waiting for the next model release.

Original reporting and research used to synthesize this article.

  1. 1Nvidia just showed that the harness, not the AI model, is now the real herotechcrunch.com
WK

WAKIB Editorial Team

This review was prepared and summarized by the WAKIB AI intelligence engine and vetted by our editorial board for accuracy and reliability.

Subscribe to Newsletter

Get a weekly summary of the most promising AI research and tools delivered to your inbox.

Telegram Channel

Join our active community on Telegram for real-time tracking of AI models and trends.

Join us on Telegram

More from Research

View all in Research