Adding Benchmaxxer Repellant to the Open ASR Leaderboard

To fight test-set contamination, the Open ASR Leaderboard adds private evaluation datasets from Appen and DataoceanAI, keeping the default average WER on public sets while exposing optional private metrics.

Wednesday May 6, 2026 Source: huggingface.co
TL;DR — Quick Answer

On May 6, 2026, the Hugging Face Open ASR Leaderboard added 11 private English ASR evaluation datasets curated with Appen and DataoceanAI, covering scripted and conversational speech across Australian, Canadian, Indian, American, and British accents. The private sets stay hidden to stop benchmaxxing — model developers gaming leaderboard scores by training on test data. The default average WER still uses only public datasets, and users can toggle the private data on to see a Rank Delta, while no per-split scores are published to avoid single-provider gaming.

Key Takeaways

Adding Benchmaxxer Repellant to the Open ASR Leaderboard — AI news article illustration

Hugging Face’s Open ASR Leaderboard — visited over 710,000 times since it launched in September 2023 — is adding what the team calls “benchmaxxer repellant.” Working with Appen and DataoceanAI, it will evaluate models on 11 private ASR datasets that only the maintainers can see.

Why Private Data

The leaderboard prizes standardization and openness — all test sets are consolidated, outputs are normalized, and the UI and evaluation scripts are open source. Openness also breeds a failure mode: benchmaxxing, where models get better on the leaderboard without getting better in the real world, whether by training on public test sets or on data resembling them. Private sets make gaming far harder.

The New Datasets

Because there is no single catch-all ASR model, the additions expose gaps between saturated settings — scripted American English — and harder conditions.

A Toggle, Not a Mandate

The default Average WER remains computed only on public datasets, so rankings do not shift underneath the community. A new Private data tab lets users optionally include the private sets in the macroaverage and watch a Rank Delta column show how ordering changes.

Safeguards Against Gaming

What’s Next

Hugging Face is exploring real-world noisy-condition evaluations, having built tooling to flag low signal-to-noise clips and transcript mismatches. With community-suggested data channels and new aggregate metrics, the leaderboard is evolving from a single number into a nuanced map of where speech recognition actually breaks.

Frequently Asked Questions

What is benchmaxxing?

Benchmaxxing is when model developers improve leaderboard scores by specifically optimizing against public test sets or training on data that closely resembles them, without real gains in robustness. It follows Goodhart's law: when a measure becomes a target, it ceases to be a good measure.

How does the Open ASR Leaderboard stop cheating?

Hugging Face added 11 private ASR datasets from Appen and DataoceanAI that are never released to developers. The default average WER excludes the private sets, users can toggle them on optionally, and no per-split scores are published so developers cannot game a single accent or provider.

What does the private data tab measure?

It reports an Average WER plus aggregate scripted, conversational, US, and non-US macroaverages computed across the data providers, so accent and condition gaps stay visible even though individual split scores are hidden.

Can I still add my model to the leaderboard?

Yes. You open a pull request on the leaderboard GitHub repo with results on the public sets; maintainers verify public results and compute the private metrics for you.

This article is based on the official announcement from huggingface.co . Read the original for full technical details.

Related Articles

Back to all news