The Open Agent Leaderboard

An open benchmark from IBM Research that compares full AI agent systems — not just models — across six tasks, reporting both quality and cost per task.

Monday May 18, 2026 Source: huggingface.co
TL;DR — Quick Answer

The Open Agent Leaderboard, launched May 18, 2026 by IBM Research, is an open benchmark that evaluates complete AI agent systems — planning, tools, memory, error recovery, and model together — across six established benchmarks instead of scoring models in isolation. It reports both average success rate and cost per task for every agent-plus-model configuration. Early results show the same model paired with different agent wrappers produces different scores and costs, failed runs cost 20 to 54 percent more than successful ones, and open-weight agents trail frontier closed-source models by 18 to 29 percentage points on average.

Key Takeaways

The Open Agent Leaderboard — AI news article illustration

The Open Agent Leaderboard, launched by IBM Research on May 18, 2026, aims to answer a deceptively hard question: how good are general-purpose AI agents? Instead of scoring models in isolation, it evaluates the full agent system — the model plus its tools, planning, memory, and error recovery — across six established benchmarks, and reports both quality and cost.

Measuring Systems, Not Just Models

Most AI evaluations report one thing: what score each model got on which task. But when you deploy an agent, you are choosing a whole system — which tools it can use, how it plans, what it remembers between actions, and how it recovers when something goes wrong. Change any of those and the same model produces very different results at very different costs.

The leaderboard treats the full agent system as the thing being measured. That makes visible what is actually driving results: gains that come from the model versus gains that come from the agent design, and which components generalize across settings.

The Six Benchmarks

The evaluation assembles six research-vetted benchmarks that test very different kinds of work:

A unified task, context, and actions protocol gives every benchmark the same shape, so each agent speaks one language while keeping its native tools.

What the Results Already Show

The top three rows of the leaderboard all use the same model, yet they differ in both score and cost — proof that the agent wrapper matters as much as the base model. Failed runs also behave very differently across agents: in the team’s experiments, failed runs cost 20 to 54 percent more than successful ones, so failure behavior shapes the bill as much as success does.

Tool shortlisting — helping an agent focus on relevant tools — improved performance across every model tested, turning otherwise failing configurations into viable ones. Since launch, two open-weight models (DeepSeek V3.2 and Kimi K2.5) were added, bringing the table to five models across five agents and six benchmarks; open-weight agents are competitive on specific combinations but trail closed frontier models by 18 to 29 percentage points on average.

What’s Public Today

Everything behind the leaderboard is open from day one:

What’s Next

The team is inviting contributions on three axes: new agents wrapped in the Exgentic protocol, new benchmarks with programmatic evaluators, and new models, especially open-weight ones. Results are submitted by opening a PR on the open-agent-leaderboard results dataset. General-purpose agents are too important to be evaluated behind closed doors — the hope is that this becomes a shared standard for the whole community.

Frequently Asked Questions

What is the Open Agent Leaderboard?

The Open Agent Leaderboard is an open benchmark from IBM Research that evaluates complete AI agent systems — the model plus its tools, planning, memory, and recovery logic — across six established benchmarks instead of scoring models in isolation.

How is the Open Agent Leaderboard different from model benchmarks?

Most benchmarks report a score per model. The leaderboard treats the whole agent system as the unit of measurement and reports both quality and cost, so the same model paired with different agents can rank very differently.

Which benchmarks does the Open Agent Leaderboard use?

It combines SWE-Bench Verified, BrowseComp+, AppWorld, tau2-Bench Airline and Retail, and tau2-Bench Telecom, unified by a shared task-context-actions protocol while keeping each agent's native tools and interfaces.

Is the Open Agent Leaderboard open source?

Yes. The leaderboard Space, the Exgentic evaluation framework on GitHub, and the methodology paper on arXiv are all public, and anyone can submit results by opening a PR on the results dataset.

This article is based on the official announcement from huggingface.co . Read the original for full technical details.

Related Articles

Back to all news