Introducing the Ettin Reranker Family

Hugging Face releases six CrossEncoder rerankers built on the Ettin ModernBERT encoders, from 17M to 1B parameters — state-of-the-art at their sizes on MTEB retrieval and up to 8.3x faster than prior baselines.

Tuesday May 19, 2026 Source: huggingface.co
TL;DR — Quick Answer

Hugging Face released six new Sentence Transformers CrossEncoder rerankers on May 19, 2026, from 17M to 1B parameters, each state-of-the-art at its size on MTEB and NanoBEIR retrieval benchmarks. The 17M model tops the speed charts at 7,517 pairs per second, and the 1B model matches its 1.54B teacher within 0.0001 on MTEB while running 2.4x faster. All six ship under Apache 2.0 and accept up to 8K tokens of context.

Key Takeaways

Introducing the Ettin Reranker Family — AI news article illustration

Hugging Face has released the Ettin Reranker family — six new CrossEncoder rerankers from 17M to 1B parameters that are state-of-the-art at their respective sizes on retrieval benchmarks. Each is built on the Ettin encoder suite from Johns Hopkins University and ships under the Apache 2.0 license with its full training recipe.

The Release

The headline: the 17M — the smallest model beats the 33M ms-marco-MiniLM-L12-v2 by +0.051 NDCG@10 on MTEB at roughly half the parameters, and the 68M model nearly matches the 596M Qwen3-Reranker-0.6B using a ninth of the compute.

What a Reranker Does

A cross-encoder takes a (query, document) pair and lets both texts attend to each other through every transformer layer — far more accurate than separate embedding vectors, but too expensive to run over a whole corpus. The standard pattern is retrieve-then-rerank: a fast embedder pulls the top-K candidates, then the reranker re-orders just those K. These models make that final re-ranks stage both cheaper and better.

Architecture

All six share a modular head: a Transformer module with sequence unpadding so Flash Attention 2 never computes on padding tokens, CLS pooling, and a two-dense-layer scoring head. Ablations showed CLS pooling edges out mean pooling despite ModernBERT’s local-window attention in most layers.

Speed

The 17M model processes 7,517 pairs per second on an H100 — nearly double the throughput of the MiniLM-L6-v2 it outperforms. The 150M model hits 3,237 pairs per second, a 2.3x gap over the two peer 150M ModernBERT rerankers that load through padded AutoModel paths. Combined bfloat16 plus unpadded FA2 delivers speedups from 1.71x on the 17M to 8.26x on the 1B.

The 1B Flagship

The 1B model scores 0.6114 on MTEB(eng, v2) Retrieval — within 0.0001 of its 1.54B teacher mxbai-rerank-large-v2 — while running at 928 pairs per second to the teacher’s 387. Distilling from a stronger teacher than the current one is, the author notes, the likely path to closing the remaining gap to Qwen3-Reranker-4B.

What This Means

Small, fast, open rerankers quietly decide the quality of every search, RAG pipeline, and recommendation system on the web. A 17M model that beats a 2x-larger legacy incumbent costs next to nothing to self-host — and with Apache 2.0 weights and a published recipe, this line of work makes best-in-class ranking a commodity.

Frequently Asked Questions

What is a reranker?

A reranker, also called a cross-encoder, scores each query-document pair jointly and re-orders the top candidates a fast embedding model retrieves, so the final ranking is both cheap and accurate.

Which Ettin reranker should I use?

For most retrieve-then-rerank pipelines the 17M or 32M models are drop-in upgrades over legacy MiniLM rerankers — higher accuracy at higher speed. The 400M and 1B versions deliver near-teacher quality when compute budget is not the constraint.

How fast are the Ettin rerankers?

On an H100 with bfloat16 and Flash Attention 2, throughput ranges from 7,517 pairs per second for the 17M model down to 928 for the 1B model — up to 8.3x faster than fp32 baselines.

Are the Ettin rerankers open source?

Yes, all six models, the ~143M-triple training dataset, and the complete training recipe are released under the Apache 2.0 license on Hugging Face.

This article is based on the official announcement from huggingface.co . Read the original for full technical details.

Related Articles

Back to all news