Hugging Face’s Open ASR Leaderboard — visited over 710,000 times since it launched in September 2023 — is adding what the team calls “benchmaxxer repellant.” Working with Appen and DataoceanAI, it will evaluate models on 11 private ASR datasets that only the maintainers can see.
Why Private Data
The leaderboard prizes standardization and openness — all test sets are consolidated, outputs are normalized, and the UI and evaluation scripts are open source. Openness also breeds a failure mode: benchmaxxing, where models get better on the leaderboard without getting better in the real world, whether by training on public test sets or on data resembling them. Private sets make gaming far harder.
The New Datasets
- Appen: scripted Australian, Canadian, Indian, and American accents plus conversational Indian and two American sets
- DataoceanAI: scripted American and British, plus conversational American and British totaling roughly 8.8 and 6 hours respectively
- All transcriptions punctuated and cased, with conversational sets preserving disfluencies
Because there is no single catch-all ASR model, the additions expose gaps between saturated settings — scripted American English — and harder conditions.
A Toggle, Not a Mandate
The default Average WER remains computed only on public datasets, so rankings do not shift underneath the community. A new Private data tab lets users optionally include the private sets in the macroaverage and watch a Rank Delta column show how ordering changes.
Safeguards Against Gaming
- Datasets are withheld from developers, and Appen and DataoceanAI were asked not to sell them to clients
- No per-split scores are published, so nobody can optimize for one provider or accent
- Multiple providers balance the advantage any single one could confer
What’s Next
Hugging Face is exploring real-world noisy-condition evaluations, having built tooling to flag low signal-to-noise clips and transcript mismatches. With community-suggested data channels and new aggregate metrics, the leaderboard is evolving from a single number into a nuanced map of where speech recognition actually breaks.