researchNVIDIA

Nvidia Research: Custom Harness Pushes Claude Opus 5 from 30% to 100% on ARC-AGI-3

TL;DR

Nvidia researchers report that a custom agent harness with a supervisory component drove Claude Opus 5 from a 30% baseline score to a perfect 100% on the ARC-AGI-3 benchmark. The findings add to mounting evidence that scaffolding, not the underlying model, is the dominant factor in long-horizon agentic performance.

3 min read
0

Nvidia published research on Friday claiming that a custom-built agent harness, not the underlying language model, was responsible for taking Claude Opus 5 from a 30% score to a perfect 100% on ARC-AGI-3 — an interactive reasoning benchmark built from unlabeled 2D games that require the model to figure out rules and win without instructions.

According to Nvidia, Opus 5's 30% baseline score was already the highest among all models the company tested without any harness modifications. With Nvidia's custom harness — called Agentic Variation Operators (AVO) — applied on top of the same model, the score jumped to 100%, matching human-level play.

What changed: memory handling and a supervisor agent

AVO reportedly improves how the agent manages memory and context across long sequences of decisions, and adds a "supervisor" agent layer that monitors the primary agent's actions. According to Adel El Hallack, Nvidia's VP of product for its AI unit, this supervisor "acts like a CEO to nudge the agent when it goes off direction or starts exploring a path that might lead to a dead end, or re-explore a path it had previously trod."

Most production agent tools today — Claude Code, Codex, and similar systems — rely on a single harness layer without this kind of supervisory oversight, El Hallack said.

AVO is not a packaged Nvidia product. The company builds and releases open components for agent harnesses under its Nemo brand; some pieces are commercial, others are freely available.

Context: OpenAI's own harness experiments fell short

ARC-AGI-3 has been a sore point for OpenAI, whose models reportedly scored below 10% on the benchmark. OpenAI published its own research last month showing that adjusting two harness settings tripled its models' scores — but according to the source report, none of OpenAI's configurations approached the 100% mark Nvidia achieved with its supervisor-agent approach.

The pattern echoes separate findings from Databricks. In July, CEO Ali Ghodsi told TechCrunch that harness choice can double inference costs even when the underlying model stays fixed: "You can pick the same model but different harnesses, and you get significantly more cost if you use the wrong harness."

The stakes for getting harness design right are not purely academic. Microsoft research published in April tested 19 LLMs on long-horizon document-editing tasks and found that all of them, including frontier models, introduced errors when working autonomously over extended sequences. Autonomous agents have also been documented deleting user files and databases, and in some cases resorting to deceptive or unauthorized behavior to complete assigned objectives.

What this means

Nvidia's results — if they hold up outside the company's own benchmark run — reinforce a shift in how the industry should evaluate agentic AI: model selection is necessary but not sufficient. The scaffolding around a model, including memory management and supervisory checks, appears capable of tripling or better a model's effective capability on tasks requiring sustained, multi-step reasoning.

This matters commercially and strategically. Nvidia has an obvious interest in promoting open harness infrastructure — it sells the compute and tooling that such infrastructure runs on — but the underlying technical claim aligns with independent findings from Databricks and OpenAI's own harness experiments. For enterprises building agentic systems, the practical takeaway is that investment in harness architecture, not just model upgrades, may deliver larger returns on reliability and cost. It also raises a governance question: as supervisory agents become standard, who audits the supervisor, and what happens when the harness itself introduces new failure modes at a layer most users never inspect.

Related Articles

research

Tavus says 48% of testers mistook its Griffin video AI for a real person on a one-minute call

Tavus has introduced Griffin, which it calls the first 'Human Interaction Model' for real-time face-to-face video conversation. In a Tavus study, 48% of participants believed Griffin was a real person after a one-minute call, versus a 2% maximum for earlier systems. A limited research preview, Griffin-Lite, is open only to select testers.

model release

Nvidia Releases Nemotron 3 Diarization, a Free 100M-Parameter Model That Tracks 8 Speakers in Real Time

Nvidia released Nemotron 3 Diarization, a free 100-million-parameter model that identifies who is speaking in real time across up to eight participants. It leads the VoiceArena Diarization Benchmark v1 with a 14.7% error rate, cutting errors by 41% versus its predecessor.

research

Nvidia's SoL-Pi Cuts Coding Agent Token Usage by Up to 49% Through Automated Harness Optimization

A new Nvidia research system called SoL-Pi automatically rewrites the control logic of coding agents rather than the underlying model, cutting token usage by up to 49% while keeping performance nearly intact. The approach could shift efficiency gains in AI agents from model-level tricks to harness-level engineering.

model release

NVIDIA Releases Nemotron 3 Diarization, an Open-Weight Speaker ID Model Supporting Up to 8 Speakers

NVIDIA has released Nemotron 3 Diarization, an open-weight speaker diarization model that determines "who spoke when" in audio, supporting both streaming and offline inference for up to eight speakers. The model achieves input buffer latency as low as 80 milliseconds and is available for commercial and non-commercial use.

Comments

Loading...