Nvidia Harness Boosts AI Reasoning to 100 Percent Score
Nvidia researchers used a custom software harness to enable the Claude Opus 5 model to achieve a perfect score on the ARC-AGI-3 reasoning benchmark.
Researchers at Nvidia demonstrated that the software scaffolding surrounding an AI model, referred to as a harness, is more critical for long-horizon tasks than the underlying model. Using a custom harness called Agentic Variation Operators (AVO), which incorporates a supervisor component to nudge the agent, the team enabled the Claude Opus 5 model to achieve a 100% score on the ARC-AGI-3 interactive reasoning benchmark.
This performance represents a significant leap over the model's 30% score without the harness. The results also surpass previous attempts by OpenAI, whose models initially scored below 10% on the same benchmark before harness adjustments tripled those figures. Nvidia argues that an open agent stack—granting control over the harness, infrastructure, and runtime—is essential for improving accuracy and security, contrasting this with the closed approaches of competitors.
Parallel research from Databricks indicates that these software choices also impact financial efficiency. The company found that the specific harness selected can significantly alter operational costs, with some configurations potentially doubling the expense of running the same AI model.