IBM Research launches Open Agent Leaderboard, showing same models achieve different results based on agent architecture
IBM Research has launched the Open Agent Leaderboard, the first open benchmark that evaluates complete AI agent systems rather than just underlying models. The leaderboard reveals that agents using identical models can achieve significantly different success rates and costs depending on system architecture, with failed runs costing 20-54% more than successful ones.
IBM Research launches Open Agent Leaderboard for complete AI agent systems
IBM Research has released the Open Agent Leaderboard, the first open benchmark designed to evaluate full AI agent systems rather than just the models that power them. The leaderboard reports both quality and cost metrics across six diverse benchmarks, revealing that agent architecture significantly impacts performance even when using identical models.
Testing generality, not specialization
The benchmark evaluates agents across six established tasks spanning different domains: SWE-Bench Verified (code bug fixes), BrowseComp+ (web research), AppWorld (personal task completion), and three tau2-Bench variants (customer service and technical support). According to IBM Research, agents are tested as general-purpose systems without benchmark-specific tuning.
The leaderboard shows that the top three configurations all use the same underlying model but achieve different success rates and costs due to variations in agent architecture. "Same model, different agents, different results — the agent matters," the researchers write.
Failed runs cost 20-54% more
One of the most significant findings concerns failure behavior. IBM Research reports that failed agent runs cost 20-54% more than successful ones, with some agents failing fast and cheaply while others burn through expensive runs before terminating. This cost differential matters for production deployments where failure patterns directly impact operational expenses.
Tool shortlisting improves all models
The research identifies specific architectural components that improve performance. Tool shortlisting, which helps agents focus on relevant tools rather than searching through all available options, improved performance across every model tested. The technique "turned otherwise failing configurations into viable ones," according to the paper.
General agents match specialized systems
Contrary to expectations, IBM Research found that general-purpose agents are already competitive with specialized ones. "Across most benchmarks, general agents match or even outperform the best specialized systems," the researchers report. This suggests that single agents can increasingly handle diverse tasks without job-specific customization.
Unified protocol enables comparison
The technical foundation is Exgentic, an open evaluation framework that implements a unified protocol. The protocol standardizes how different agent systems interact with benchmarks by providing a consistent structure: a task description, context information, and available actions. This standardization allows fair comparison across agent architectures while preserving each benchmark's original design.
What this means
By evaluating complete agent systems rather than isolated models, the leaderboard makes visible what drives real-world performance: planning strategies, memory management, tool selection, and error recovery. The finding that failed runs cost significantly more than successful ones highlights a critical operational consideration typically absent from model benchmarks. Most importantly, the open release of methodology, framework, and results enables the research community to reproduce evaluations and submit new agent configurations. This transparency is essential for understanding which architectural choices generalize across tasks and which improvements come from the model versus the agent wrapper. For teams deploying AI agents, the leaderboard provides the first standardized way to compare full system costs and capabilities rather than relying solely on model performance claims.
Related Articles
Aleph Alpha benchmark: Chinese AI models balanced on just 17-41% of 967 sensitive-topic prompts
An Aleph Alpha study of 967 politically sensitive prompts found that only 17 to 41 percent of responses from Alibaba's Qwen, DeepSeek and Moonshot's Kimi were balanced, according to the company's own AI scorer. DeepSeek V4 Pro refused roughly two-thirds of questions. Aleph Alpha sells "sovereign AI" to governments, which gives it a commercial interest in the result.
Mercor study: AI models beat 12 licensed CPAs on simplified accounting tasks, but top score on full APEX benchmark is 61
A Mercor study found AI models beat 12 licensed CPAs on simplified tasks from the APEX Accounting Benchmark. On the full 160-task benchmark, the top model, Claude Opus 5.5, meets 61.8% of grading criteria, and no model fully solved almost 60% of tasks.
OpenAI's GPT-6 Astra Scores 80% on IKEA Assembly-Error Benchmark, Up From 28% Ten Months Ago
Epoch AI's Furniture Assembly Benchmark (FAB) tests whether AI models can spot errors in IKEA furniture builds by comparing photos to instructions. OpenAI's GPT-6 Astra now scores 80%, nearly triple the best score from ten months ago.
Robot Safety Benchmark Finds GPT-6 Astra and Claude Fable 5.1 Rarely Refuse Dangerous Commands
A new benchmark called RoboHarm tested whether AI models controlling robotic arms would refuse dangerous commands. GPT-6 Astra completed 60 of 100 dangerous tasks and Claude Fable 5.1 completed 34, with neither model showing a reliable safety layer.
Comments
Loading...