The Open Agent Leaderboard, launched by IBM Research on May 18, 2026, aims to answer a deceptively hard question: how good are general-purpose AI agents? Instead of scoring models in isolation, it evaluates the full agent system — the model plus its tools, planning, memory, and error recovery — across six established benchmarks, and reports both quality and cost.
Measuring Systems, Not Just Models
Most AI evaluations report one thing: what score each model got on which task. But when you deploy an agent, you are choosing a whole system — which tools it can use, how it plans, what it remembers between actions, and how it recovers when something goes wrong. Change any of those and the same model produces very different results at very different costs.
The leaderboard treats the full agent system as the thing being measured. That makes visible what is actually driving results: gains that come from the model versus gains that come from the agent design, and which components generalize across settings.
The Six Benchmarks
The evaluation assembles six research-vetted benchmarks that test very different kinds of work:
- SWE-Bench Verified — fixing real bugs in real code repositories
- BrowseComp+ — researching complex questions across the web
- AppWorld — completing personal tasks across hundreds of apps and actions
- tau2-Bench Airline and Retail — customer service following company policies
- tau2-Bench Telecom — technical support under policy constraints
A unified task, context, and actions protocol gives every benchmark the same shape, so each agent speaks one language while keeping its native tools.
What the Results Already Show
The top three rows of the leaderboard all use the same model, yet they differ in both score and cost — proof that the agent wrapper matters as much as the base model. Failed runs also behave very differently across agents: in the team’s experiments, failed runs cost 20 to 54 percent more than successful ones, so failure behavior shapes the bill as much as success does.
Tool shortlisting — helping an agent focus on relevant tools — improved performance across every model tested, turning otherwise failing configurations into viable ones. Since launch, two open-weight models (DeepSeek V3.2 and Kimi K2.5) were added, bringing the table to five models across five agents and six benchmarks; open-weight agents are competitive on specific combinations but trail closed frontier models by 18 to 29 percentage points on average.
What’s Public Today
Everything behind the leaderboard is open from day one:
- The Open Agent Leaderboard Space for exploring results
- Exgentic, a framework for running and reproducing cross-environment evaluations
- The paper, covering full methodology and empirical analysis
What’s Next
The team is inviting contributions on three axes: new agents wrapped in the Exgentic protocol, new benchmarks with programmatic evaluators, and new models, especially open-weight ones. Results are submitted by opening a PR on the open-agent-leaderboard results dataset. General-purpose agents are too important to be evaluated behind closed doors — the hope is that this becomes a shared standard for the whole community.